tech, developers, and the code underneath

issue 021· news·

GTC 2025: a roadmap to 2027 and a warning about power

Blackwell Ultra this year, Rubin next, Feynman after. Nvidia is now publishing a schedule like a foundry.

Jensen Huang did two hours at the San Jose arena and the substance came down to one slide: a named product cadence stretching to 2027.

  • Blackwell Ultra (GB300) — second half of 2025, more HBM per package.
  • Vera Rubin — 2026, new CPU (Vera) and new GPU (Rubin), HBM4.
  • Rubin Ultra — 2027, a much larger rack-scale configuration.
  • Feynman — 2028, named, not detailed.

Publishing a multi-year roadmap at this granularity is a foundry move, not a product-company move. The audience is not developers. It is the utilities, the construction firms, the HBM suppliers, and the CFOs who need to plan capital around it.

the number that should worry you#

The power figures for the rack-scale systems are the part of this keynote that will still matter in five years. Current NVL72 racks draw on the order of 120 kW. The Rubin Ultra generation is being discussed in the hundreds of kilowatts per rack.

A traditional enterprise datacenter rack is provisioned for 5 to 15 kW. Air cooling tops out somewhere around 40 kW with heroic effort. Everything past that is liquid, and everything past about 150 kW is liquid plus a fundamentally different power distribution architecture — which is why the roadmap includes 800 VDC distribution.

Translation: existing datacenter shells are largely unusable for this. The buildout is not "install new servers," it is "build new buildings near new substations." That is a multi-year, capital-intensive, permit-bound process, and it is the actual rate limiter on AI capacity through the rest of the decade.

the software announcements#

Dynamo, an open-source inference serving framework, is the developer-relevant release. It handles disaggregated serving — running the prefill phase and the decode phase on different hardware pools, because they have completely different compute and memory characteristics. Prefill is compute-bound and parallel; decode is memory-bandwidth-bound and sequential. Running them on the same homogeneous pool wastes a lot of silicon.

If you operate inference at any scale, disaggregation is probably the largest single efficiency win available to you right now, and having a maintained open implementation lowers the bar considerably.

NIM microservices continue to be Nvidia's attempt to own the deployment layer. Reasonable, containerized, and a lock-in vector you should evaluate with clear eyes.

the strategic read#

Nvidia's actual product is no longer a chip. It is a rack, plus the networking between racks, plus the software stack on top. Selling a GPU means competing with AMD and with every hyperscaler's internal silicon team. Selling an integrated rack-scale system with a co-designed interconnect means competing with nobody, because nobody else has the interconnect.

That is the moat, it is being widened deliberately, and the roadmap slide is the announcement of that strategy rather than a product list.

Dom, March 19, 2025

get README in your inbox

One dispatch, no noise. Tech and developer news, plus the occasional long piece on the craft.

subscribe →