The graphics processor that powers the AI boom is, by ancestry, a machine for drawing triangles: a flexible engine that got repurposed, brilliantly, into the workhorse of deep learning. That flexibility was indispensable for research, where the shape of the problem keeps changing. But most of the compute in the industry is no longer research; it is serving, which means running finished models billions of times a day to answer queries. And there, generality is a tax. A chip that can do anything spends silicon, power and cost keeping options open that a deployed model never uses. The quiet correction underway is a fleet of inference ASICs: chips stripped down to the specific arithmetic of running a fixed network, faster and cheaper per answer precisely because they can do almost nothing else. Training will keep its flexible engines. Serving is discovering that the profitable move is to build hardware that has forgotten how to be anything but a model.
Your free preview ends here
Keep reading with a free account
Sign up and redeem two free articles per month. No credit card required.
Already a member? Log in

