The AI chip race is shifting from “single-GPU performance” to a more pragmatic question:How many tokens can be generated per kilowatt-hour?Nvidia’s latest disclosed data shows that on the next-generation Vera Rubin NVL72 exist DeepSeek V4 Pro running a 1.6-trillion-parameter model, token throughput per megawatt is 7× that of Blackwell.
Nvidia’s New Accounting Methodology
Specific figures: Under an assumption of 100 tokens per second per user, Rubin delivers 59.4 million tokens/second per megawatt, versus GB300’s 28.5 million. Nvidia has also introduced a new metric—“Profit per gigawatt-year”—with Rubin at approximately $14.99 billion and GB300 at approximately $10.53 billion. This clearly targets the industry pain point that “power supply and capacity are the bottlenecks,” as power consumption is precisely the wall every buyer is hitting this month.
Note thatVera Rubin Vera Rubin is Nvidia’s next-generation architecture following Blackwell (named after astronomer Vera Rubin), and the “7×” figure is a modeled projection—not a measured result. Yet even halving it would sufficiently explain why all hyperscale customers are rushing orders for Rubin before Blackwell has fully depreciated—power allocations are finite, and higher output per unit of energy consumption means more users served within the same data center footprint.
Meta Accelerates In-House Chip Development
Meanwhile, Meta has confirmed its in-house MTIA 450 chip will enter mass production in the first half of 2027MTIA 500 Follow-up by end of 2027: the former doubles HBM bandwidth over the previous generation, while the latter boosts it by another 50% and increases HBM capacity by up to 80%, targeting approximately 44% TCO (total cost of ownership) savings versus GPUs, aligned with a roughly $115 billion capital expenditure plan.
One commonality is:Memory has become central to all in-house chips. Meta’s HBM addition and Positron’s launch last week of 2304 GB per die both point to the same conclusion—what’s currently scarce is not compute, but memory.
Another disruptor
Also noteworthy is Cornelis: this company raised $205 million to build a GPU-agnostic interconnect layer, competing with InfiniBand and NVLink; its 400 Gbps CN5000 is already shipping, and its 800 Gbps CN6000 will ship in Q4. Its significance lies in enabling buyers to mix Nvidia, AMD, and in-house chips within the same cluster.
For ordinary users, these changes will ultimately manifest in two ways: continued decline in AI service costs and longer uptime. To stay updated on AI tools and infrastructure developments, visit AI Dash.
