Nvidia Vera Rubin Achieves 7× the Energy Efficiency of Blackwell; Meta’s In-House Chip Enters Mass Production in 2027

The AI chip race is shifting from “single-GPU performance” to a more pragmatic question:How many tokens can be generated per kilowatt-hour?Nvidia’s latest disclosed data shows that on the next-generation Vera Rubin NVL72 exist DeepSeek V4 Pro running a 1.6-trillion-parameter model, token throughput per megawatt is 7× that of Blackwell.

Nvidia’s New Accounting Methodology

Specific figures: Under an assumption of 100 tokens per second per user, Rubin delivers 59.4 million tokens/second per megawatt, versus GB300’s 28.5 million. Nvidia has also introduced a new metric—“Profit per gigawatt-year”—with Rubin at approximately $14.99 billion and GB300 at approximately $10.53 billion. This clearly targets the industry pain point that “power supply and capacity are the bottlenecks,” as power consumption is precisely the wall every buyer is hitting this month.

Note thatVera Rubin Vera Rubin is Nvidia’s next-generation architecture following Blackwell (named after astronomer Vera Rubin), and the “7×” figure is a modeled projection—not a measured result. Yet even halving it would sufficiently explain why all hyperscale customers are rushing orders for Rubin before Blackwell has fully depreciated—power allocations are finite, and higher output per unit of energy consumption means more users served within the same data center footprint.

Meta Accelerates In-House Chip Development

Meanwhile, Meta has confirmed its in-house MTIA 450 chip will enter mass production in the first half of 2027MTIA 500 Follow-up by end of 2027: the former doubles HBM bandwidth over the previous generation, while the latter boosts it by another 50% and increases HBM capacity by up to 80%, targeting approximately 44% TCO (total cost of ownership) savings versus GPUs, aligned with a roughly $115 billion capital expenditure plan.

One commonality is:Memory has become central to all in-house chips. Meta’s HBM addition and Positron’s launch last week of 2304 GB per die both point to the same conclusion—what’s currently scarce is not compute, but memory.

Another disruptor

Also noteworthy is Cornelis: this company raised $205 million to build a GPU-agnostic interconnect layer, competing with InfiniBand and NVLink; its 400 Gbps CN5000 is already shipping, and its 800 Gbps CN6000 will ship in Q4. Its significance lies in enabling buyers to mix Nvidia, AMD, and in-house chips within the same cluster.

For ordinary users, these changes will ultimately manifest in two ways: continued decline in AI service costs and longer uptime. To stay updated on AI tools and infrastructure developments, visit AI Dash.

🔗 Share: Twitter Weibo Copy link

📬 Like this article?

Weekly selected AI tool reviews + practical tutorials, delivered directly to you.

Subscribe to the weekly AI picks →

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top
Tool Picks
1
AI Writing
GPT-6.1 Sol Deep Review: OpenAI’s efficiency model evolves again—five times cheaper, performance approaching Astra
8.8
📊AI Productivity 💻AI Coding 📝AI Writing 🎨AI Image Gen
📬 Weekly AI Picks