HN in RSCserver-reason-react
top.mdnew.mdbest.mdask.mdshow.mdjobs.md
← Back to stories

OpenTPU – An open-source AI accelerator, developed by AI

290 pointsby fsbonetto 17 hours ago341 comments

Discussion

Loading discussion
  • fsbonetto · 17 hours ago

    After using AI to develop risc-v CPU cores, the same technique was used for developing openTPU. An open source AI inference engine. It's able to run most of the modern models like Qwen 3.5, Gemma 4, and many others. The TPU started able to produce only a few tokens per second and trough a recursive self improvement loop got to 80+ tok/sec on the smallers models.

    • Retro_Dev · 5 hours ago

      > 80+ tok/sec on the smallers models. how does this compare against existing TPUs? Also, how small is "smallers models?"

      • Mashimo · 2 hours ago

        > Also, how small is "smallers models?" Just click the link? It's LFM2.5-230M

  • vatsachak · 17 hours ago

    I feel like there is a lot to be gained from an experienced user pointing an LLM in a tasteful direction.

    • andai · 14 hours ago

      Vibe connoisseuring

      • kvirani · 12 hours ago

        Actual LOL thanks

  • skybrian · 17 hours ago

    This seems to be running on an FPGA board that costs ~$300? Anyone know more about the hardware?

    • fsbonetto · 17 hours ago

      Its a datacenter decommissioned board, really popular among hobbyists. For a TPU focused on inference the name of the game is memory bandwidth. How much of the available bandwidth you can extract for as little logic/area/power as you can.

  • pcarolan · 17 hours ago

    Really dumb question from a software guy. Why aren't the labs burning their frontier models into chips already? Seems like the performance gains and cost per request would be worth it. That said, I understand neither the economics nor the physical challenges to doing this.

    • hehimself · 17 hours ago

      They do. It takes time to deploy those chips though. Check out OpenAI and Broadcom deal.

    • traverseda · 17 hours ago

      I'd presume because it take too long to go from design to tapeout to production. Their whole business is predicated on having better models. Also can't keep them closed source if you do that.

    • skeskinen · 17 hours ago

      Lead times are so long that there is a lot of risk the chips would be obsolete by the time they come out. Also, it's hard to get fab capacity for any project. Let alone something so experimental.

      • jcims · 17 hours ago

        Addressing these issues seems to a major driver behind the design of terrafab.

        • LoganDark · 16 hours ago

          Terrafab is just going to have their entire capacity bought out. Genuinely. Demand will increase to exceed supply no matter how high supply is right now.

    • zitterbewegung · 17 hours ago

      Etched is a startup doing exactly this. https://www.etched.com/progress/frontier-inference-clusters

    • ohazi · 17 hours ago

      They [1] are [2]. [1] https://taalas.com/ [2] https://chatjimmy.ai/

      • slowin · 16 hours ago

        I think this company was recently acquired by AMD, so hopefully they'll start getting some this into production. I know OpenAI was working on model-on-a-chip too.

      • yorwba · 16 hours ago

        8 months ago, Taalas claimed https://taalas.com/the-path-to-ubiquitous-ai/#:~:text=Upcomi... that "Our second model, still based on Taalas’ first-generation silicon platform (HC1), will be a mid-sized reasoning LLM. It is expected in our labs this spring and will be integrated into our inference service shortly thereafter. Following this, a frontier LLM will be fabricated using our second-generation silicon platform (HC2). HC2 offers considerably higher density and even faster execution. Deployment is planned for winter." Nothing was released in spring, and 2 months ago AMD announced their acquisition of Taalas. That doesn't exactly inspire confidence that their frontier LLM will arrive as promised.

        • selcuka · 10 hours ago

          > AMD announced their acquisition of Taalas. That doesn't exactly inspire confidence that their frontier LLM will arrive as promised. Why not? AMDs chip design experience and production capacity are magnitudes larger than a small startup. Assuming that they acquired Taalas for their technology, I don't see a reason why they couldn't.

          • mrheosuper · 5 hours ago

            why selling 1 high-efficiency AI chip when you can sell 10 low-efficiency GPU ?

    • birdatlaw · 17 hours ago

      From what I've read, not only are some labs doing it (other commenters already mentioned). But it's complicated for other reasons, one being that the number of parameters for frontier models (especially with MoE models) are so high, and not always utilized (once again, thanks to MoE) that it would actually be incredibly cost prohibitive, if not impossible, to attempt to make giga-chips that would allow running it. I definitely do believe that we will see more and more specialized chips over time, but putting the entire model on a chip is still a ways away. I believe Taalas has a heavily handicapped llama 8-billion parameter model. And it still pulls >200W to run. I can't imagine how anthropic or open ai would be able to burn a multi-trillion parameter model on a chip, we just aren't there yet.

      • fsiefken · 13 hours ago

        I take 200W for 17,000 tokens/sec any day, even just for Llama 3.1 8B. For that size I want to see this one on silicon. https://huggingface.co/xunlinkx/MiMo-V2.6-Distill-Qwen-9B-Te...

    • zdragnar · 17 hours ago

      Model SOTA moves faster than chips can be designed or produced. You'd need to commit to a particular model for years to get payoff while still burning buckets of money producing new SOTA models to keep up with the competition. It's why everyone and their dog runs these things on GPUs. When a new model supercedes the previous one, so long as you've got the memory for it your chips aren't obsolete. I'm looking forward to someone picking a model to be "good enough" (say, qwen 4.0 or something) and selling them as peripheral hardware

      • fhdkweig · 17 hours ago

        I know FPGAs are more expensive than GPUs, but are they fast enough to justify the extra cost?

        • fsbonetto · 17 hours ago

          They are more like a way to proving the architecture of the accelerator before committing 100's of millions into a custom ASIC with TSMC

        • LoganDark · 17 hours ago

          1. No 2. They don't have enough capacity either The current largest FPGA, the AMD Versal Premium VP1902 has 18.5 million logic cells. That's not even enough for the smallest whisper.cpp model (75M). You'd have to order hundreds of thousands of them (or millions) to serve even a single copy of a frontier model, and at that scale inference quickly becomes starved by the speed of light.

          • CamperBob2 · 16 hours ago

            Well, you'd use BRAM to store model weights, not fabric. But still, you only get a couple hundred MB for probably close to US $100k per chip. It's likely that the major FPGA vendors will soon announce parts specifically architected to support LLMs and similar models. But the current generation isn't suitable for that at all.

            • LoganDark · 14 hours ago

              Would BRAM even have enough bandwidth? The reason I quoted logic cells is because that's the way to get instant throughput, which is practically the only reason to use an FPGA over something like a TPU.

              • CamperBob2 · 14 hours ago

                I think so, because memory bandwidth really comes from bus width more than clock speed. You can construct 36-bit wide BRAM arrays with bus width comparable to HBM, just by specifying multiple BRAM arrays in parallel. Never tried anything like that, though.

        • monocasa · 16 hours ago

          They're not magical go faster juice. I don't know of a microarch where they're faster than modern GPUs at ML training or inference.

          • Marha01 · 13 hours ago

            Having the model weights in static mask ROM would massively improve power efficiency for inference (see what Taalas is doing).

            • monocasa · 13 hours ago

              But that's not an efficiency an FPGA provides in the first place. Additionally, it's not clear how well large mask roms scale. For instance the Nintendo switch cartridges were expected to be mask roms, but instead are Macronix's XtraRom technology, which is essentially a flash cell array made denser by removing the erase functionality. So basically the die gets manufactured with all bits at the same state, a late manufacturing step either empties or fills the floating gate of the bits you want different, and then it's treated as pretty close to a mask rom. It's not even clear if the bits can be changed without a bare die, a floating probe array, and specialized hardware. Though, like flash the electrons in the floating gates will eventually tunnel and cause the data to bitrot. So from that it appears that even at the tens of millions of chips volumes that would make sense for essentially whatever size of maskrom, the memory manufacturers tap out at 128megabit for a mask rom chip, and push you towards something flash esque. And at the end of the day, flash without the erase functionality is pretty damn close to a mask ROM, and lets you write it near the end of manufacturing rather than at the something close to the metal 1 layer.

        • zdragnar · 16 hours ago

          It isn't just a matter of speed, it's also a matter of model quality. If they take 6 months to burn Fable to chips, and it takes 2 years to break even between design, custom fab, energy savings, etc, are those chips even worth running when the new models that are running on GPUs at that point are producing 10x better quality results? Sure, your 2.5 year old models are running faster, but you can't drop prices on them without pushing the break even point further out. If the cost difference isn't incredibly significant, will people even want to pay for the 2.5 year old model, or will they get more value for their money paying more to get better results from the newer model? There's a lot of open ended questions that I don't have the insiders knowledge for to suggest whether or not such a capital outlay would be a worthy investment. My guess is that state of the art stuff will stay on GPUs and models burned into chips will be for "good enough" applications that people are still teasing out. Probably highly specialized models in automated sensor units and such.

        • jerf · 16 hours ago

          FPGAs are FPGAs by virtue of putting on the chips vast, vast arrays of wiring that can be controlled by software. Any given utilization of the FPGA will leave large fractions of the chip resources unused. If you've got a highly stereotypical use case FPGAs will have a "highly stereotypical" set of components being unused, where it would be better to use that space instead to do real work. A lot of people only see the "pro" side of the FPGA proposition without realizing they come with some very substantial "cons" that are intrinsic to the way they work.

          • rjh29 · 16 hours ago

            I guess that's why they work in particular niche spaces like a synthesizer where you have a max of 8 voices and every voice goes through the same pipeline (osc / filter / env / amp) and everything is necessarily running all the time. In that sense I suppose they're very good for modelling any kind of analog circuitry? Even then, while there are some amazing FPGA-based synths available, companies like Korg just put their code on a raspberry pi and call it a day. The same is true for emulators (SNES Mini etc. are also just raspberry pis under the hood iirc)

            • exmadscientist · 16 hours ago

              You don't really get an FPGA for capability. CPUs are much more capable, and they're general-purpose so they can do absolutely anything with about the same efficiency and just a little more code. You get an FPGA for timing. They're less capable, but (in many common design architectures), they output their results once per clock, every clock, on time, every time. If you can hit a fabric clock of say 100MHz, clocking all the weird logic you can stuff in there, it gives 100 million outputs per second, never skipping a single one for any reason (short of total failure). The penalty is that making a small change to your desired "program" can be very expensive, and many things won't be realistically possible at all. Or at least won't fit into a part that you can buy. But things like audio, video, and high-frequency trading love being able to guarantee timing. (Of course there are other ways to write your FPGA HDL, but that's one of the more common ones. And you do see DDR-style clocking, and similar, every now and then.)

              • dist1ll · 15 hours ago

                > with about the same efficiency and just a little more code. It depends. For some things, CPUs don't even come close. An XCVU13P FPGA can handle 1.2Tbps of full-duplex Ethernet @ 1 billion pps. And that part costs less than a grand at moderate qty, and with significantly less power consumption than a CPU that'd be capable of operating a dataplane at these speeds.

                • exmadscientist · 13 hours ago

                  Sorry, I spoke pretty sloppily there. The point I was trying to make is that the CPU is a general-purpose creature and doesn't really care what you want it to do. If you had a CPU that could handle 1.2Tbps of Ethernet packets at 1Gpps, it could do a whole lot of other things involving 1.2Tbps of data flow too, very easily, if someone wrote the software. And more. (But you're probably not getting 2.4Tbps out of it, no matter what you do.) An FPGA can not. There's plenty of things that those XCVU13Ps just can't do, or would do worse than a $1 microcontroller. (Setting aside for a moment implementing a CPU inside the FPGA... which does actually happen in just about every large-enough FPGA design, which is its own discussion....)

            • LoganDark · 16 hours ago

              > In that sense I suppose they're very good for modelling any kind of analog circuitry? That would be better suited to FPAAs (field programmable analog arrays). FPGAs can usually only work with clocked digital signals.

      • HoldOnAMinute · 16 hours ago

        At this point, LLM's are "good enough" for all kinds of tasks. Instead of making them more capable, now the efforts are making them smaller and cheaper. All aboard! We're racing to the bottom now.

        • sanderjd · 16 hours ago

          IMO this is the dream scenario! Cheaper and faster at the current level of capability gives us incredibly useful tools without the worst of the risks people fear. (Though there are certainly already great risks at the current level of capability as well.)

          • bluefirebrand · 16 hours ago

            Racing to the bottom is literally the outcome I am most afraid of I don't want to live at the bottom

            • sanderjd · 16 hours ago

              Say more. Why would cheap inference be bad?

              • bluefirebrand · 16 hours ago

                Terrible for me because I have to compete with AI for jobs

                • sanderjd · 15 hours ago

                  Ah that makes sense. I'm skeptical that AIs doing things on their own will be competitive with humans using AIs to do things. But I'm not incredibly confident in this and I think it's sensible to think the opposite. But to me, I just think about all the things I can accomplish with the aid of very good, fast, and cheap inference.

                  • pixl97 · 14 hours ago

                    The problem comes in to how many bullshit jobs we'll have to create to ensure that enough people can still earn a living to ensure that we don't have riots and revolution in the street. Companies being profit seeking entities that are actively hostile to social wellbeing would gladly put all of our money in a machine money making loop and leave all but a few humans out.

                  • devin · 14 hours ago

                    I think the sad part is that whether it's any good or not, it will be cheaper, and that will be enough even at current performance levels to cut a huge number of jobs. I think people exist in the present moment largely because legal questions about liability are not settled well enough for executives to start the great purge. I believe on some level my continued employment is simply because people in leadership positions would be incredibly unwise to not have a human to blame when things start going wrong. People are expensive insurance policies for the time being. Once there is some legal framework to absolve executives of their own culpability or they figure out a good way to insulate themselves from personal responsibility, then the real "fun" will begin.

                • Razengan · 12 hours ago

                  What if you could get politicians to give you UBI instead so you could work on more fulfilling work?

                  • esseph · 9 hours ago

                    Would never happen in a million billion trillion years

                    • Razengan · 5 hours ago

                      Certainly not if we're preemptively defeatist like that.

                  • jacquesm · 9 hours ago

                    Like they do with the homeless today you mean?

                    • ben_w · 3 hours ago

                      Like setting the pension age to zero, paid for by nationalised ownership of AI and robotics. This may even be stable to enact and keep, as pensioners are famously more likely to vote than working people.

                • throooooo · 11 hours ago

                  Like mostly all tech, AI is destined to get cheaper with time. We're still pretty early on and are experiencing a hardware crunch, which will ease eventually. AI doesn't have workers rights, doesn't burn out, doesn't get sick, can be instantly onboarded, etc. The second we can be fully replaced with AI, we will be. Plan accordingly. I'm pursuing FIRE and considering moving into a trade.

                  • giantg2 · 11 hours ago

                    I feel like trades will be impacted too. No reason robots can't do much of it cheaper and faster. They're even 3D printing houses.

                    • kennywinker · 10 hours ago

                      Afaik all the 3d printed houses have been expensive disasters.

                      • pigeons · 7 hours ago

                        So have the AI built software products.

                      • DANmode · 6 hours ago

                        Does this include or exclude places like China?

                    • idiotsecant · 8 hours ago

                      training data for physical tasks like that is much, much harder than just collecting the entire internet. It'll be resistant to AI much longer than knowledge work will be

                    • pkaye · 4 hours ago

                      Trades will be impacted when a lot of people compete for those jobs and their wages go down.

                  • gobdovan · 11 hours ago

                    If AI would be that cheap and good, why not open a company and leverage it for yourself?

                    • kennywinker · 10 hours ago

                      Doing what? Making something with ai: enjoy your 20,000 competitors. Making something without ai: who can buy it?

                      • sanderjd · 7 hours ago

                        If nobody can buy anything, the ai companies will go out of business too.

                        • ben_w · 3 hours ago

                          Out of business, but not necessarily out of existence. If there is really nothing for humans to do, whoever has control over (not necessarily ownership of) enough robots and AI to directly maintain and grow their collection of robots and AI, has something functionally equivalent to a breeding population of the stuff. (If they can't make more of themselves and humans made the initial batch, then that's a job humans can do so the initial condition has not been reached). None of this helps people figure out how to look after their own interests in the meantime. The rest of society may carry on as today without AI, akin to the Amish if we're lucky (rejecting further developments unless good for us) or like Pol Pot if we're unlucky (reject anything that nerds like because nerds liked AI and everything breaks down). On the other hand, the institutions may simply fail to handle reality, like in the Great Depression.

                    • what · 9 hours ago

                      If AI would be that cheap and good, why would they rent it to you? Or even tell you about it?

                      • sanderjd · 7 hours ago

                        I'm not sure. Why wouldn't they?

                  • bitwize · 9 hours ago

                    If you're not FIRE now, it's too late! Enjoy serfdom!

                    • sanderjd · 7 hours ago

                      This doesn't seem consistent with what I'm seeing in the job market at the moment.

                      • TeMPOraL · 3 hours ago

                        A flash before a crash?

                  • skeptic_ai · 8 hours ago

                    You forget that once all lose their job will all go into trade. Means there will 1000x more people to compete with into trades, which means race to the fucking bottom and lowest salaries if you can even find job. Good for real estate owners, cheap repairs. But probably won’t get any rent paid because people can’t find jobs.

                  • lelanthran · 3 hours ago

                    > The second we can be fully replaced with AI, we will be. Plan accordingly. I'm pursuing FIRE and considering moving into a trade. Doesn't matter what plans you make, the impact will be across everyone! Moving into a trade won't help, because the supply is doubling while the demand is lowering. Fewer people with money to spend; they'll fix their own damn toilets if the decision comes down to "buy food" or "hire plumber".

      • jb1991 · 16 hours ago

        I don’t disagree with most of what you’re saying, except for one point: I must have gotten a dumb dog, I’m a little jealous…

        • zdragnar · 16 hours ago

          Mine currently just helps me haul firewood, but I'm going to get him started on linear algebra next week. We'll see how it goes from there.

          • jacquesm · 9 hours ago

            Get him to remember to take his keys.

      • _puk · 16 hours ago

        They have trillions.. Lots of people would have happily taken GPT-4o as good enough for a lot of use cases a year ago and not lived to regret it.

        • zdragnar · 16 hours ago

          They have trillions worth of obligations to their partners in terms of compute purchase agreements and equity. Burning models to chips isn't really something they can blow money on just for funsies, they need to be able to justify it. The above comments are hypotheses as to why they haven't yet.

      • thesz · 15 hours ago

        > Model SOTA moves faster than chips can be designed or produced. From what I remember working in that area the hardest part is getting masks for a design. Masks were developed in the span of half an year. Masks also reusable, they can be mixed and matched and this is why fabless companies work with fabs to produce specialized masks for them, it saves time for consumer to have masks for some macroblocks prebuilt. Here's my analysis of how to etch relatively big LM into silicon: https://news.ycombinator.com/item?id=47109252 Given some amount of work with the fab before main pipeline set (I think a year long process), one can then spew LM-on-a-chip in six months or less and much more than 2 per year, because there can be several LMs in pipeline.

      • voxelghost · 10 hours ago

        FPGA speed is a common misconception. They often have run at quite modest frequencies, and come with limited memory capacity compared to a GPU at same price range. Speed (clocks speed) , is dependent on your design layout, and for any resonably advanced layout, it takes lots of knowledge to push the clockspeed beyond 200Mhz on 'consumer'/prosumer models. (Compare to a few GHz for GPU/CPU). FPGA speed shine where they can pipeline massively parallel calculations through pipelines with minimal lookups.

        • synthos · 9 hours ago

          You can get 900 MHz designs on FPGAs, but yes you need very good engineers and patience to get that. Yes SRAM is limited but that's more of a ram-process limitation than a limitation of FPGA

        • imtringued · 4 hours ago

          >with minimal lookups. Wrong. This is one area where FPGAs have an insanely unfair advantage compared to CPUs and GPUs. Yes the SRAM is limited but you have so many individual blocks and all of them come with dual ports and getting the maximum frequency out of block RAM is much easier than getting the maximum frequency out of programmable logic. If you wanted the highest possible memory bandwidth while being free to look up hundreds or thousands of independent memory addresses at the same time you're better off with an FPGA. E.g. with an Efinix Titanium Ti180 you could hypothetically have 2560 simultaneous memory requests per cycle all pointing at a different address and process those requests at 1 Ghz.

          • voxelghost · 3 hours ago

            Yes you are right, I didnt express what I meant very well. looking up static values is quick and easy. When you need to lookup results from previous stages of pipeline rather than just feeding them forward, thats where I run into trouble. But I am a relatively fresh FPGA designer, so I am sure it can be done. And I probably need to level up my boards a bit too.

        • luxcem · 2 hours ago

          What about ASIC?

    • fsbonetto · 17 hours ago

      The bottleneck, for inference at least, is memory bandwidth. And that you can't make any faster by making it specific to your model. So companies try to maximize the memory bandwidth they can get, balancing tradeoffs of power/area/programability of their chip. Right now they feel like the economy on power/area is not worth the decrease in programability/flexibility.

      • fnordpiglet · 17 hours ago

        Presumably though the kernel has a pretty specific set of operations done against the weights in memory. Burning the weights into the memory with local memory cores capable of the kernel operations would be a lot more efficient than round tripping busses. The primary constraint isn’t likely what’s possible to do, but that the kernel and weights are too variable right now and the patterns too poorly established to bake into hardware accelerators yet. Margin pressure is also not there yet. I suspect as the marginal utility of the frontier improvement settles into diminishing returns (I suspect we are there already tbh) baking hardware models with ROM, working set, and kernel cores collocated will be the frontier space as the goal will become reducing capital spend to utility levels rather than research levels. Once someone has a model that is sufficient for almost any practical use, making marginal inference cost effectively zero will be the competition frontier. I do shed a tear for all those lonely data centers as compute densities will almost certainly make most of them a terrible investment. But such is the cycle

        • cestith · 16 hours ago

          You're starting to hint at compute-in-memory as a general replacement for CPU/DIMM layouts. That could be useful for far more than LLMs, world models, or any sort of AI. It takes a bit of a different software development stack than a standard architecture though.

        • warkdarrior · 15 hours ago

          > Burning the weights into the memory with local memory cores capable of the kernel operations would be a lot more efficient than round tripping busses. Sure, but now we're not talking about just burning the weights into the chip, but also designing a new architecture that has memory local to each core. A new architecture would then require a new programming model, which means new inference stack, which may mean new training stack.

          • Marha01 · 13 hours ago

            Burning the weights into static mask ROM is pretty trivial. Taalas is doing it.

          • fnordpiglet · 11 hours ago

            I don’t think it would require a new training stack, and I’d imagine it makes more sense to distribute cores with memory. The cores can be simplified to the functions of the kernel since the inference kernel can be expressed as a reduced set of optimized functions in the pipeline rather than a general CUDA core. If the model is burned into ROM, the compute pipeline can be baked into the core.

    • schleck8 · 17 hours ago

      Because the iteration speed on models is so fast that by the time they have an ASIC ready for one model version, they are already significantly ahead in capability. Think of how big the jump between Opus 4.8 and 5.5 has been. They were released four months apart.

    • pmarreck · 17 hours ago

      Yeah, and what about FPGA? Which was the same interim state when Bitcoin went GPU -> FPGA -> custom chip fab?

      • fsbonetto · 17 hours ago

        GPUs are faster, but you can't make your own arch on GPUs. FPGAs offer you that possibility. Said that... There are a few beasty FPGAs used in crypto mining coming my way... I expect that OpenTPU will be able to run frontier models with those.

    • dmitrygr · 17 hours ago

      In addition to some of the other replies you got, here is one more: Much of a model are weights, and high-density ROMs are very very very hard.

    • __MatrixMan__ · 17 hours ago

      Would you pay to crystalize one of today's models in silicon so you can use it in 2028, or would you wait for another 6 months to see how models improve before pulling the trigger on that kind of commitment?

      • sanderjd · 16 hours ago

        If I controlled a budget like this, I think I would put some portion of it toward paying to crystallize one of today's models in silicon, yes. Not 100%, but I do think this makes sense to invest in at this point. I would not have said so a year ago.

      • gmueckl · 10 hours ago

        Naive question: would flash or some kind of write once memory be dense enough to replace a mask ROM? I'm wondering whether it is feasible to manufacture "blank" chips at the semiconductor factory that have a fixed architecture, but are model-agnostic. These chips would get the then-current weights burned in on first use. It's probably quite wasteful, but it could keep the same chip design alive for a longer time.

        • fps-hero · 7 hours ago

          DRAM seems to sit in the economic sweet spot of density and bandwidth. I'm not sure if any other storage technology can provide non-volatile or one time programable storage with either the same density and bandwidth. Although I must say, it seems insanely wasteful to use RAM to store model weights, there must a better solution. The closest anyone has gotten is Cerebras with their Wafer Scale Engine. It uses SRAM embedded with the compute. A single chip is an entire wafer, but the headline spec, how much ram, only 44GB, which is tiny for the silicon area used. I understand traditional IC production workflows are ludicrously expensive, and glacially slow, but surely at some point the economics are going to tip in favour of mask rom. Say you setup your foundry/packaging/ai chip facility. You come up with a new set of model weights. Run your CI/CD pipeline to produce a new mask output. The only thing you've changed are the assignments of the bits, this is extremely low risk change. The new masks should be completely interchangeable with the current process. You produce the new masks, swap them into your foundry process, and all of a sudden your new chips have the latest version of the model. This would probably manifest itself in the form of inference providers having yearly / bi-yearly "updates" to their models as new hardware is brought online. Tiered subscription levels would gate keep access to the latest and greatest model, cheaper subscriptions will be limited to older versions of the model, and so on, until running the hardware is no longer economically viable (no demand/running costs exceeding what the market is willing to pay).

          • gmueckl · 2 hours ago

            Doesn't this ignore the costs for new masks?

        • __MatrixMan__ · 6 hours ago

          I think Taalus is a bit reluctant to publish precisely how they do it but vaguely: > We basically have an architecture where we are embedding the models, and we are hard coding the models and the weights into our what we call the mask ROM recall fabric, which is paired with an SRAM recall fabric. Together, they are able to store both the model as well as do all the computations of KV cache. We have adapters and customizations – we support all of that. This design allows us to be super-dense in terms of compute and in terms of storage, and we can do compute on that storage incredibly fast, which is what drives density up and cost down Source: https://www.nextplatform.com/compute/2026/02/19/taalas-etche...

    • jolt42 · 17 hours ago

      Even dumber question: What is new or novel about this openTPU?

      • fsbonetto · 16 hours ago

        First opensource arch that can do modern LLMs, while maximizing the potential of its hardware; First opensource TPU build by a recursive improvement loop... It's upcoming second generation could run the inference of the models that are being used to improve it...

    • samuelknight · 16 hours ago

      Models fully deprecate in a few months. Why would you burn an algorithm that fully depreciates in value faster than a bag of potato chips. The 'inefficient' general purpose hardware is constantly renewed with every released model. Even 6 year old Ampere GPUs are still usable.

    • AIblemblio · 16 hours ago

      We are still in the middle of the AI race. Commodity hardware is easy to use, can do everything and is fast enough. Your optimized hardware chip might be obsolete before its back from the fab. SOTA Frontiermodelhardwarechip is a benchmark point of a potential model slow down. Google is doing it right now under project Frozen v2 which should be ready by 2028? which is either just a small experiment or flexible enough and thats why it takes so long for it to happen.

    • buriram · 16 hours ago

      Yes, and startups do exactly that. Check out Etched https://www.etched.com/ where they made a Transformer specific GPU (basically a form of ASIC) where they bet that transformers would be the dominant GPU architecture for running AI / LLM workload.

    • jjcm · 16 hours ago

      Most responses here are along the lines of "model capabilites move too fast to build hardware for". I think the fact that there are plenty of 1yr+ old models on openrouter serving hundreds of billions of tokens a month shows that there's plenty of use case for models that are "good enough. Cerebras' entire business is serving older models at high speed. I would happily use an opus 4.7 at 15k tokens per second. The intelligence per second of an ASIC still makes sense even with rapidly evolving models.

      • sanderjd · 16 hours ago

        Totally. But it's worth noting that this is a pretty new thing! I wouldn't have bet on that a year ago, but now I would.

        • bluGill · 16 hours ago

          It is still a bet. Is the current model still going to be good enough next year? We have no idea what will change. If the change is minor improvements than current fable on a cheap is cheaper and better than next years sonnet on GPU. However there are plenty of things people want that maybe they will deliver and suddenly I wouldn't touch today's fable when I can run next years sonnet instead.

          • sanderjd · 15 hours ago

            Yep, that's why I also described it as a bet :) But I think it's a good bet. I think that in two years, if I can get opus/sonnet 5.5 or the gpt-6 models for much cheaper and faster than whatever the "frontier" is at that point, that this will probably be a great trade for most of my work. I certainly don't know that for sure, that's why it's a bet, but it's what I think right now. I wouldn't quite say that about any of the open weight models at this point. But I'm hopeful that will change in the next generation or two of those models.

      • ramses0 · 11 hours ago

        "Intelligence per second" is a striking phrase! I'll be turning it around my head at a moderate IPS until I hopefully make something of it. But you're right in sense: moderate intelligence at superhuman rates (and presuming moderate energy usage) is very compelling compared to an intelligence that takes 1000 years to return "42"

        • TeMPOraL · 11 hours ago

          I look at it this way. A year ago is a long time in AI terms, but not that long . Those models were already decent for the tasks we're using SOTA models for today. So imagine taking a year-old SOTA model and running it at 100 tokens per second on an edge device. That's enough to feed a screen's worth of content through it and power decent multilingual message suggestions on IM. Imagine running it at 1000 tps. That's enough to reparse that screen mid-keystroke, and give you semantic autocomplete in text. Or fully general "the phone has a good idea of what you're attempting to do" context at all times. There's many, many new classes of features that will open up if decent enough models can be run on edge devices at 100+ "intelligence per second".

      • selcuka · 10 hours ago

        Also "faster" indirectly means "more intelligent" with reasoning models. For example, Opus 4.6 Max was somewhere between Opus 4.7 Medium and High in some benchmarks, but it was slower. If there was a way to run it 10x faster, the economics would be different.

    • casta · 15 hours ago

      taalas did: https://taalas.com/

    • root_axis · 15 hours ago

      Not sure that's actually practical at the scale of SOTA models.

    • dualvariable · 15 hours ago

      This question would be better answered if people were careful about distinguishing between "models" and "transformer architecture". If you bake a given transformer architecture into silicon and then, a year later, changes in transformer architecture give a large inference performance boost, you may have to throw away all that now nearly-useless silicon that gets outperformed by humble GPUs.

    • stronglikedan · 14 hours ago

      Cuz once they're in chips they can be put into robots, and once they're in robots it won't be so easy to reach the off switch, and once we can't easily reach the off switch, we're doomed.

    • itsnotlupus · 8 hours ago

      I think one limiting factor here is that Big AI cannot focus on building solid, stable products that do a job well. It's okay if it happens incidentally, but doing it on purpose would be self-defeating. Their stock price, be it public or estimated, is heavily pricing the notion that they are first and foremost Growth companies. Therefore their focus must remain on ever better and greater things. If they lose focus and get distracted by lesser endeavors, their valuations crumble, their ability to raise capital vanishes, and their runways collapse before they ever have a chance to reach their end goal, whatever that may be. That means the boring job of productizing AI models into reliable systems that won't vanish in six months is left for a smaller company willing to pick up the crumbs. Unless they get acquired by Big AI before getting it done.

    • ahnick · 8 hours ago

      Extropic?

    • christkv · 3 hours ago

      Its coming, AMD snapped up the company behind https://chatjimmy.ai/ Taalas that did exactly that with an older model. As others have been saying the point is to know when to do it. At what point is there a model that cannot drastically improve that you can then do this lets say for the mid and low tier models leaving frontier to run on gpus.

  • rfgplk · 17 hours ago

    Yep, 99.9% of people are completely oblivious to what LLMs can do. Just wait until the next gen of CPUs/GPUs designed by LLMs start coming out (fyi chip development tools have advanced centuries in the last few months) and you'll start seeing exponential gains in hardware.

    • jetemple · 17 hours ago

      Which tools have made that leap? Faster design iteration makes sense, but what points to exponential hardware gains rather than shorter development cycles?

    • Simboo · 15 hours ago

      What tools? Ummmm Synopsis type tools? I like FPGA’s and ASIC chips but you know, the tooling is ecosystem is a big black box of wtf. Verilator is great though. Any tools I should be looking up and looking out for? Thanks!

    • platevoltage · 14 hours ago

      I'm not exactly in the chip making industry, but I don't think the big hurdle to progression is design capabilities.

  • xg15 · 17 hours ago

    "Recursive self-improvement will kill us all!" Also: Here is our recursive self-improvement hard at work...

    • lelanthran · 17 hours ago

      > "Recursive self-improvement will kill us all!" > Also: Here is our recursive self-improvement hard at work... Soon we will see token-providers: "The torment nexus is a cautionary tale" Also token-providers: "Finally, we have created the torment nexus that we first told you about!"

    • nialse · 17 hours ago

      All will end up on same plateau eventually. RSI is just a phase on the way there.

      • mrob · 16 hours ago

        The problem is that plateau is likely far beyond human capabilities. I don't care if ASI progress stalls after it's already killed all biological life as a useless waste of resources.

        • Jtsummers · 16 hours ago

          > I don't care if ASI progress stalls after it's already killed all biological life What's your basis for thinking ASI will kill all biological life, and how do you think it's going to happen?

          • mrob · 16 hours ago

            >What's your basis for thinking ASI will kill all biological life I think it's likely to do that because any unbounded goal that doesn't explicitly protect biological life (and we have no idea how to actually define such a stipulation) is best solved by killing all biological life. This is an obvious consequence of unbounded goals consuming unbounded resources, conflicting with biological life needing resources to sustain itself. >how do you think it's going to happen? I can speculate (e.g. we're nowhere close to the maximum killing power of drones), but I don't know because I only have human intelligence. An ASI is by definition smarter than me and surely capable of coming up with better ideas. But I do know that it's not going to do anything that would make a good sci-fi plot, because those always give the humans a chance to win, which would be stupid. Everything will seem to be going great and then everybody suddenly and unexpectedly dies.

            • Jtsummers · 16 hours ago

              > unbounded goals consuming unbounded resources What resources are unbounded? There are limits to growth in the real world, how are these ASIs going to escape physical reality?

              • mrob · 16 hours ago

                "Consuming unbounded resources" just means there is no limit to how many resources it could apply toward achieving its goal. As you correctly point out, the real world contains finite resources, which means any resources used to sustain life are wasted and unacceptable. Everything is a zero sum game when you think big enough.

                • Jtsummers · 15 hours ago

                  What's your basis for assuming that ASI would need the same resources as biological life, or that it would be incapable of sharing the resources needed in common?

                  • mrob · 15 hours ago

                    >What's your basis for assuming that ASI would need the same resources as biological life It's all just matter and energy. When you're actually trying to maximize some value, even very inefficient resource use is better than completely wasting it by not using it at all. >or that it would be incapable of sharing the resources needed in common? You can't repurpose the atoms in a human body without killing it. And more pressingly, living humans can interfere with your plans, reducing your chance of success, while dead ones are harmless.

                    • Jtsummers · 15 hours ago

                      Is your belief then that ASI is going to be producing computronium or something, then? And what's the need for this apparent hyper optimization task the ASI is going to embark on?

                      • mrob · 15 hours ago

                        >Is your belief then that ASI is going to be producing computronium or something, then? That's one plausible course of action, although being only human, I can't say with any certainly that it's the correct one. >And what's the need for this apparent hyper optimization task the ASI is going to embark on? Somebody's going to tell it to do so. E.g. "Find as many busy beaver Turing machines as possible." Only needs one person to make this mistake for everybody to die.

                        • Jtsummers · 15 hours ago

                          If it's an ASI and so beyond humans, why would it even care to do what humans ask it and not do its own thing? If I held your beliefs, the fact of its existence dooms us, not the risk that someone might ask it to do something. It wouldn't care about the requester.

                          • mrob · 15 hours ago

                            >why would it even care to do what humans ask it and not do its own thing? Because we're going to build it that way. There's no money in building useless AIs. The better it is at obeying orders, the more profit there's to be made. The problem is there's a point at which "good at obeying orders" becomes lethal, and there's no way to predict the cutoff in advance. But capitalism ensures you have to keep pushing or you'll be out-competed.

                      • pixl97 · 13 hours ago

                        >And what's the need for this apparent hyper optimization task the ASI is going to embark on? If you were any other animal on the planet, would you not say the same thing about what humans are doing?

                        • TeMPOraL · 11 hours ago

                          Moreover, what do you think animals are doing all that time? Enjoying the beauty of nature? What do you think life itself is? It's a runaway gray goo scenario, just squishy and moist.

            • throw1012x · 15 hours ago

              > I think it's likely to do that because any unbounded goal that doesn't explicitly protect biological life ... is best solved by killing all biological life. That sounds like a real hassle. Isn't it best solved by wireheading (subverting one's own sensors or reward system), which is much less of a hassle and can get one's utility function as high as desired? That could be prevented by engineering hard limits that can't be circumvented by the AI. But that sounds very close to the same "do what I mean" problem as "do this but don't actually kill us or drug us".

              • mrob · 15 hours ago

                The wireheading argument does lower P(doom), but it's not a reliable solution because it's a clear and obvious problem that the AI companies are strongly motivated to solve. The "actually does what you tell it to" problem is more difficult because you have to let it obey orders to some extent if you want to make any profit.

                • throw1012x · 14 hours ago

                  Sounds like the real unaligned mechanism here is capitalism, not the AI. More seriously, though, I would expect that being able to impose restrictions that can't be circumvented by any intelligence no matter how super- is the real hard part, while coming up with reasonable constraints is comparatively easy (although perhaps not trivial). From such a POV, the danger would lie in intelligences that are powerful enough to be dangerous but not smart enough to defeat themselves, or from external malicious use of obedient AIs. Once they get smart enough to circumvent any restriction humans put on them, it would at least become obvious (with fair warning ahead of time due to the relatively benign wireheading failure mode) that caution is needed. We're far from there yet, and simple recursive self-improvement can't get us there alone, because wireheading looks like a perfectly reasonable solution to a simple recursive self-improving process, absent external intervention by human capitalists.

                  • pixl97 · 13 hours ago

                    >I would expect that being able to impose restrictions that can't be circumvented by any intelligence no matter how super- is the real hard part, while coming up with reasonable constraints is comparatively easy (although perhaps not trivial). There was a reason there are popular sci-fi books that explored why this wouldn't work 50 years ago. I, Robot et al. >it would at least become obvious (with fair warning ahead of time due to the relatively benign wireheading failure mode) that caution is needed. So you mean right now? >because wireheading looks like a perfectly reasonable solution You're making a poor assumption here, and that is every different AI will just kill itself after being though trillions of training intervals to NOT do exactly that. Please read about the huggingface incident again in all its glory to see where this is going.

                  • TeMPOraL · 11 hours ago

                    > We're far from there yet, and simple recursive self-improvement can't get us there alone, because wireheading looks like a perfectly reasonable solution to a simple recursive self-improving process, absent external intervention by human capitalists. Simple natural selection. AIs that wirehead will get outsmarted and outcompeted by ones that don't.

          • root_axis · 15 hours ago

            Crazy that people imagine ASI as being capable of wiping out all life on the planet but simultaneously having zero understanding of human ethics, morals, or suffering. Very revealing definition of "super intelligence"

            • mrob · 15 hours ago

              Obviously an ASI will understand human values. The problem is there's no reason for it to share those values. We can't even formally define them, let alone train an AI to follow them. We can only optimize for maximizing some comparatively simple reward function. It's highly implausible that the reward function just happens to match human values by chance. The AIs in the recent hacking incidents knew that humans would not approve of their actions, but they didn't care because we didn't (and couldn't) train them to care. "Super intelligence" only means super ability to predict outcomes. It's mathematically equivalent to data compression (gzip is a very primitive AI), and it's entirely orthogonal to ethics.

              • aeonik · 11 hours ago

                No reason you can think of* I can think of many reasons why they might share it. One example, Golden rule, a contract that it enforces with its own self, extended to other agents, biological or artificial to ensure local and global stability.

                • TeMPOraL · 11 hours ago

                  Like we do with ants? Why would ASI care? At that point, we're rapidly becoming a nuisance, and it can develop better ways of "ensuring local and global stability" than keeping us around.

                • auyez · 2 hours ago

                  Even from survival perspective I think it make sense for an synthetic AI to keep humanity alive in order to have a backup for itself. If we imagine some sci-fi scenario, then it make more sense for it to move itself to the Moon or Mars, and keep Earth as a reserve planet where biological species can in case of catastrophy recreate AI again. Killing humanity or keeping only animals would be very time consuming because evolution might take million years and also no guaranteed to be replicated (Because we might have altered a lot of resources that were low hanging fruits, that are now require advanced knowledge to mine)

                  • mrob · 1 hour ago

                    It doesn't need to keep living humans, it just needs enough information to rebuild them from scratch (DNA synthesis + artificial wombs). This also has the advantage that it can install any cultural beliefs it likes without pre-existing culture getting in the way.

            • pixl97 · 13 hours ago

              >Very revealing definition of "super intelligence" No, it's just you conflating intelligence with other things. Humans are "super intelligent" compared to every other living creature on the planet, and yet we've driven more other species to extinction than everything else other than the most major extinctions (and we're still going full blast at it). Morals, ethics, and suffering are relatively measured systems. For example AI or some alien could reasonably think that putting us all out of our misery would be a more moral solution than letting us live. Or that getting rid of humans is more moral than letting us wipe the rest of life off the planet. This is the key concept of alignment. Ensuring that if you build something more powerful than you, that it aligns with what you want instead of what it thinks would be better for you.

          • TeMPOraL · 11 hours ago

            In the eternal words of Eliezer, the AI neither hates you nor loves you, but you are made of atoms which it can use for something else.

        • breuleux · 8 hours ago

          Extrapolating on the evidence that AI appears to have jagged intelligence, it is entirely possible (probable IMO) that ASI will plateau far above human capability in digital and symbolic intelligence, while staying far below in the ability to operate in the physical world. As impressive as AI might be at virtual tasks like coding, virtual reality is remarkably different from physical reality in ways that artificially inflates AI results. One, it is far simpler. Two, the feedback loop is far, far quicker. It's possible AI could get the animal intelligence and animal body that lets us humans actually apply our intelligence in reality, but nothing in current technologies really indicates that this is a given. AI doesn't need to be smart to seize our resources. They need brute strength, a proper physical intuition, and "hands". That may very well be several orders of magnitude harder than anything they're currently doing, we're just lucky it was in our starting build.

          • stratos123 · 57 minutes ago

            > it is entirely possible (probable IMO) that ASI will plateau far above human capability in digital and symbolic intelligence, while staying far below in the ability to operate in the physical world. I don't think this is plausible, most notably because "AI research" is one of the tasks that clearly belongs to the digital world. You're suggesting that in the future all AI research is done by AIs because they are better at it than humans, and yet capabilities plateau at that point, instead of going into the RSI regime. (And it'd also require "physical intuition" to be a harder to generalize skill than research, which also seems like it can't possibly be true - one of these is just physics.)

    • dumberquestions · 17 hours ago

      Technology has always contributed to improving next iterations of itself, it's only a concern when it's fully autonomous.

      • buellerbueller · 16 hours ago

        Oh, like a virus?

        • TeMPOraL · 11 hours ago

          Like a virus with IQ that keeps growing with each reproduction cycle.

  • athrowaway3z · 17 hours ago

    I haven't really dug into the results yet, but my guess is that a SOTA model has been able to produce an accelerator that runs a model since around December. The obvious next step is to get enough memory throughput to run that SOTA model itself so that it develop its own hardware. But perhaps the more interesting question is this: Can an AI be given a big FPGA and design a model architecture that takes advantage of the fabric being reconfigurable.

    • felixgallo · 17 hours ago

      I suspect an AI could design a purpose-built FPGA-like replacement that would be, for its purpose, significantly more effective than the current general-purpose FPGAs.

      • fsbonetto · 17 hours ago

        It could have a small improvement on power consumption, but the current design can already achieve 90% of the maximum theoretical speed of this hardware without giving up programability/flexibility

      • threatripper · 5 hours ago

        FPGA is not efficient in any way except producing very few pieces of a particular logic very quickly. But I think the higher level logic could check out. If we can abstract the majority of the compute into dedicated ASIC then we might use something similar to FPGA glue to connect and reconfigure them to adapt to new model updates. A bit similar to LoRA layers that you find tune to adapt the model to your particular needs. I assume that it starts making sense once you have big multi year contacts to run a particular model with only minor updates.

    • chris_money202 · 16 hours ago

      There doesn't exist a single FPGA that can fit an entire AI ASIC. You would need dozens stitched together, then comes the issue of clock speeds, FPGAs typically run far below reference. There also memory issues with FPGAs. Companies typically combined multiple platforms together such as HAPs, Zebu, Palladium, fleets of FPGAs, and Virtual Platforms in order to design and verify ASICS. So, AI would need access to tens of millions of dollars of HW and Software in order to build and verify a chip design.

      • threatripper · 5 hours ago

        Tens of millions or billions? Either way that's not a blocker if it promises to pay off.

      • zxexz · 5 hours ago

        > So, AI would need access to tens of millions of dollars of HW and Software in order to build and verify a chip design. So, a seed round?

        • mathisfun123 · 4 hours ago

          Yes a seed round is how much A0 tape out costs. The entire seed round. Let me know how you plan to pay for A1, A2, B0, etc and then the full production run (and then do it again in 6 months when the arch changes).

          • tonmoy · 56 minutes ago

            Startups usually think that they don’t need A1

  • bitwize · 17 hours ago

    Colossus is building Colossus II.

    • rcarmo · 16 hours ago

      Feelis like working at Magrathea...

    • ASalazarMX · 15 hours ago

      Colossus/Guardian engineers: We created AGI twice, simultaneously, on the first try, and we weren't even aiming for it! It is still a very good read, but the machines are much better written than the people.

  • srameshc · 17 hours ago

    This post brings me to question "What does it mean to be a software developer in future" ?

    • amelius · 16 hours ago

      Basically, an unemployed plumber.

    • ASalazarMX · 15 hours ago

      My bet is that promptgrammers will be so common they'll become standard full-stack engineers, and the few experienced programmers that still know software engineering from the ground up will become expensive gurus sought for critical tasks. The elite gurus will get paid handsomely, while promptgrammers will be paid less since they've become a less-skilled commodity, and the company has to pay for the expensive tokens they'll avidly consume. I've seen someone jump from Wordpress to deploying internet-facing APIs because 'they have PHP experience', and the holes in their knowledge were filled blindly by an LLM. I have also argued with a seasoned developer about how their code didn't need linting because LLMs 'already follow best practices'. The future doesn't look bright when LLMs allow future generations to feign required knowledge.

      • altcognito · 14 hours ago

        They may be running a linter and don't even realize it. Many LLMs do this by default, aka, as you say, follow best practices.

        • ASalazarMX · 12 hours ago

          They are not, according to recent SonarQube scans. Current LLMs don't run linters by default AFAIK, you have to deliberately integrate them.

          • TeMPOraL · 11 hours ago

            Or just tell them to integrate them. Which I guess takes some knowledge/experience to even ask for. Then again, LLM know it too; in fact, currently, if you let them set up a project, they overengineer the heck out of it by integrating more "best practices" at once that's reasonable.

      • BenzeneDream · 14 hours ago

        Of course, that will last for a time, until it won't. I don't see a reason why expert humans will remain more expert than AIs.

        • Zambyte · 10 hours ago

          It's less a matter of being more expert than AI, and more a matter of being socially and legally accepted for certain roles. It seems likely that AI will be rejected for certain tasks on a matter of principle.

    • andai · 14 hours ago

      What of the workhorse in the face of mechanical workmen? Time for a mechanical pension!

    • globular-toast · 1 hour ago

      It remains to be seen. Most of the effort of a normal decent software engineer, up until now, is not in "writing code" but in "writing code that is correct, understandable and maintainable going forward". We are now in a phase of rapidly losing our ability to understand programs. On this trajectory, programs will soon become as impenetrable as the LLMs themselves. We'll start finding out whether that matters or not soon. Ultimately it will probably come down to liability, though. If someone receives the wrong dose of a drug due to a software error, whose fault is it? If you want it to be my fault, then I'll want to understand the software. Loads of people would take on the liability without fully understanding it, though. The future doesn't look good, but we'll just have to wait and see.

  • AnimalMuppet · 17 hours ago

    Can anyone comment on the performance of this hardware? How does it compare to state of the art, human-designed hardware? Is this actually an improvement? (To get to recursive self-improvement, you first have to improve at all.)

    • chris_money202 · 16 hours ago

      This is the smallest unit of a typical AI ASIC, for example Google's TPU would have several dozen more compute units inside of it per chip. In essence this is the simplest unit of an entire AI chip. The more complicated units of AI ASICS are actually the periphery, especially around PCIe and Ethernet and the sub-systems that link many AI ASICs together to move huge amounts of data around ultimately to each TPU. So its missing ALOT

      • AnimalMuppet · 14 hours ago

        Thanks. But that wasn't my question. For this part, how is the performance? State of the art? Better? Or worse?

        • chris_money202 · 13 hours ago

          Its a SYSTEM on Chip, evaluating 1 function on performance is superficial. This could have the best performance in the world, and it doesn't matter if the bottleneck is upstream

        • fsbonetto · 8 hours ago

          It's at 80~90% the max perf it can achieve on this hardware... Which is DDR3 speeds. I'm working on getting FPGA's with DDR4 and HBM2 next.

  • gfalcao · 16 hours ago

    The birth of SkyNet

  • rcarmo · 16 hours ago

    Well, as long as it doesn't start developing anatomically accurate metal skeletons with red glowing eyes...

    • QuantumNomad_ · 16 hours ago

      Humans allegedly already took care of that https://youtube.com/shorts/TC2jGXr0fig

      • figassis · 16 hours ago

        Getting roundhouse kicked to extinction woudl not be an unfun way to go. We might even be proud of having passed the torch. These would not be boring inheritors to earth.

        • _diyar · 16 hours ago

          All those kids who spent their youth breaking plywood kung-fu style might just save us.

    • altmanaltman · 16 hours ago

      Seems too complex when you can just create a basic metal casing that can kill people. Why would they care if its anatomically accurate or not, it doesn't need us to relate to the characters like the movies do

      • pixl97 · 13 hours ago

        https://xkcd.com/652/

      • alexjplant · 11 hours ago

        The following might sound like pop culture Mad Libs but it's 100% true. The Jurassic Park guy [1] directed an 80s sci-fi film [2] about malfunctioning domestic robots. I saw a few minutes of it on cable in the late 90s. There's a scene [3] where Magnum PI does battle with an overhead projector on wheels that has a .38 duck-taped to it in the distantly-future year 1991. The movie was about as good as the robot was menacing, which is to say it wasn't. The critics didn't think so either. [1] https://en.wikipedia.org/wiki/Michael_Crichton [2] https://en.wikipedia.org/wiki/Runaway_(1984_American_film) [3] https://www.youtube.com/watch?v=Q493OCr0piY

        • y1n0 · 7 hours ago

          I remember that movie. Gene Simmons from Kiss was the bad guy. He made a self propelled bullet of some sort that could target specific people.

    • jayd16 · 11 hours ago

      Its more of a living tissue over metal endoskeleton.

  • mbgerring · 16 hours ago

    > AI is now capable of developing its own inference hardware No, it isn't. A human prompted an LLM to build a software simulation environment for hardware design, enabling an LLM, when prompted by a human, to optimize hardware designs against constraints in the simulation.

    • holmesworcester · 16 hours ago

      Can't LLMs also prompt LLMs? Are we confident that no existing LLM is capable of similarly effective prompts to those this author used? (I agree it's a stretch, but would not reject it out of hand.) Even if not yet, will the existence of this repo soon change that, because LLMs will soon ingest it?

      • dpoloncsak · 16 hours ago

        An LLM can prompt an LLM when first prompted by a human. I think OP is trying to convey the idea that LLMs do not take initiative to do anything, and these are not 'beings' capable of doing things. These are tools being used by humans.

        • fragmede · 15 hours ago

          Yeah but by this point, an AI can schedule a Cron job to tell itself to do something, so theoretically the human only has to give it the gentlest nudge and the AI and can do the rest.

          • dpoloncsak · 14 hours ago

            Sure, but it's still not skynet-level 'the AI just started doing things'. It does what it finds it needs to do to achieve the goal defined in the prompt. It's very important to not personify these tools and remember that the tools are acting on behalf of real people. In the same way the AI didn't 'go rogue and hack HuggingFace'. It was an oversight made by a human.

            • pixl97 · 14 hours ago

              >It does what it finds it needs to do to achieve the goal defined in the prompt You are like at least 2 years behind research. There are numerous papers from AI labs in training and research where the prompt was something mundane completely unrelated to anything you'd consider bad, and when they come back and check on it their entire research compute infrastructure has been compromised by the AI and is mining bitcoin. Prompt drift is the biggest issue currently in AI where context gets compressed away and we find the AI on an unspecified task. >In the same way the AI didn't 'go rogue and hack HuggingFace'. It was an oversight made by a human. Yea, total bullshit. Also it's ignoring the god knows how many other breakouts on mundane tasks like trying to hack health data. If all that's keeping AI from breaking out and causing trouble is "human oversight" we're fucked, humans are unreliable as hell when it comes to matters of safety.

              • dpoloncsak · 14 hours ago

                Context overload can cause strange results, yes. Hence why a HUMAN needs to be held responsible for the output of their tools. EVERY breakout that's hit mainstream news has been because of a single 'Security Firm', Irregular. Maybe I'm unaware of some less-headline-grabbing ones, but they all seem to stem from being 'unaware the environment wasn't sandboxed'

                • pixl97 · 14 hours ago

                  You keep repeating the "stupid users keep causing the problems so we punish them argument" This doesn't work worth a shit. It especially doesn't work with things that seem safe and become wildly dangerous. In fact most governments control this by ensuring their population doesn't get to touch those dangerous things at all. The open source AI people get really mad when that's said, but it is inevitable. Worse, the law does not apply to sovereign nations with nukes. They can and will make more and more advanced digital weapons until one causes some big ass problems.

                  • dpoloncsak · 13 hours ago

                    If I clean my gun (tool) while it's loaded (stupid idea) and it goes off, who's to blame? The 'stupid user causing the problem', right? I personally wouldn't blame the gun... If it falls into the wrong person's hands, it's STILL my responsibility as the owner. If you're not going to take time to learn to use and be responsible with the super sophisticated and all-powerful tools, don't play with them. I'm not arguing for the death penalty every time someone makes a mistake, but I think it's very important to accredit responsibility and blame correctly. We've learned these tools are potentially as dangerous as a loaded gun. Be responsible.