Randomness, and the bugs you cannot reproduce
The bug that happens once a week and never in staging. A systematic approach to the class of problem everyone handles badly.
tech, developers, and the code underneath
older dispatches from the README archive.
The bug that happens once a week and never in staging. A systematic approach to the class of problem everyone handles badly.
The role is genuinely ambiguous and that ambiguity is load-bearing. An attempt at a concrete description.
Not legal advice. A working engineer's map of what the common licenses actually require you to do.
Most tracing deployments produce beautiful waterfalls nobody opens. The difference is three implementation details.
Large mechanical refactors are the clearest win available from coding agents. Here is the process that keeps them safe.
A month of synchronized global demand across three countries, sixteen cities, and every streaming platform at once.
New OS versions, more on-device model surface, and a developer relationship that remains complicated.
The tools got dramatically better while everyone was ignoring them because of a bad experience in 2015.
An instant is not a time. A time is not a date. The rules change by government decree with a few weeks' notice.
The received wisdom is never rewrite. The received wisdom is mostly right and has three real exceptions.
Six fields. Most reports have two. The difference is measured in days of engineering time.
Running code close to users is a real win for a narrow set of workloads and a complication for everything else.
Not research benchmarks. A practical harness you can build in a day that makes every future model decision an hour instead of a week.
Not the only one. But if you only internalize one, make it this one, because it makes most of the others survivable.
Internal platforms fail for predictable reasons. The successful ones share four properties.
Donations, foundations, dual licensing, open core, and hosted service. Each works for a specific shape of project.
The cheapest performance win available, ignored because it is not glamorous. Where it pays and which algorithm to pick.
Most scaling problems are solved by making something asynchronous. Most reliability problems are caused by doing it badly.
Game days, failure injection, and the specific reason your untested runbook is wrong.
The GIL-free build is officially supported and the ecosystem work is the actual story. A progress report.