tech, developers, and the code underneath

issue 012· news·

Grok 3 and the case for buying your way to the frontier

xAI's third model arrives on a cluster built in months. The interesting claim isn't the benchmark, it's the schedule.

xAI announced Grok 3 last night, along with reasoning variants and a "Big Brain" extended-thinking mode. The benchmark claims put it at or near the frontier across math, science and coding evaluations.

Take the benchmark numbers with the usual salt — they were presented by the vendor, some were shown with consensus-of-N sampling against competitors' single samples, and the field has no agreed protocol for this. The comparison charts in a launch livestream are marketing artifacts.

The genuinely notable thing is not the model. It is Colossus.

the cluster#

xAI built a datacenter in Memphis housing on the order of 100,000 H100-class GPUs, and did it on a timeline measured in months rather than years. The conventional wisdom on a buildout of that scale was eighteen to twenty-four months. They compressed it by doing things that are expensive and unglamorous: bringing in mobile gas turbines for interim power, running their own networking integration, and accepting a lot of operational risk.

Whether you find that admirable or reckless depends on your priors and on how you feel about the air quality complaints from the surrounding neighborhood, which are a real and ongoing dispute worth reading about separately.

But as an engineering datapoint it matters: it establishes that the time constant for standing up frontier-scale compute is shorter than the industry assumed, if you are willing to spend and to eat the risk.

what that implies#

If a well-capitalized new entrant can go from nothing to frontier-scale compute in about a year, then compute is not a durable moat. It is a capital requirement, which is a different thing. Capital requirements keep out the under-funded; they do not keep out the well-funded.

Which pushes the question of where the actual moat is:

  • Data — increasingly contested, increasingly litigated, and the frontier labs are all converging on similar synthetic-data-plus-RL recipes anyway.
  • Talent — mobile, expensive, and being bid on aggressively.
  • Distribution — this one is real. A model inside a product a billion people already open is worth more than a marginally better model behind a signup form.
  • Cost per token at quality — real, and derived from architecture and serving engineering rather than raw scale.

My read is that distribution and serving efficiency are the durable ones, and that is a much less romantic answer than "we have the smartest model."

for people who ship things#

Practical implication: assume model quality converges and plan accordingly. Do not build your product on the assumption that one vendor's model stays two steps ahead. Build the abstraction layer, keep your evals in your own repo, and make switching a config change.

The teams that did this in 2024 spent a boring week on it and have been changing providers casually ever since. The teams that did not are having architecture meetings.

Dom, February 18, 2025

get README in your inbox

One dispatch, no noise. Tech and developer news, plus the occasional long piece on the craft.

subscribe →