Small models ate the middle
The capability floor rose faster than the ceiling. Most production inference no longer touches a frontier model.
tech, developers, and the code underneath
20 pieces tagged models,
from January 21, 2025 to February 23, 2026.
The capability floor rose faster than the ceiling. Most production inference no longer touches a frontier model.
Twelve months that took open models from interesting to unavoidable, and where the gap actually sits.
A frontier release with a large price cut, plus effort controls and context compaction as a first-class feature.
Google ships a frontier model and Antigravity, an agent-first development environment. The bundling is the strategy.
Instant and Thinking as named modes, adaptive reasoning, and personality controls. The router lesson got learned.
A small model at frontier-adjacent coding performance, priced for volume. The economics of agent fleets just changed.
A model tuned for long-horizon autonomous work, plus checkpoints and context editing in the SDK.
One model that decides how hard to think, an abrupt deprecation of everything else, and a lesson about attachment.
gpt-oss-120b and gpt-oss-20b under Apache 2.0. A strategic reversal, six years late, and genuinely useful.
Moonshot ships a 1T-parameter MoE with 32B active, tuned for agentic tool use, with weights you can download.
xAI claims frontier results with heavy test-time compute. The number that matters is the one nobody quotes.
Pro and Flash hit GA, Flash-Lite arrives, and every tier exposes a thinking budget.