Question the measurement
My LLM judge was flipping a coin a third of the time
What happened when I tested the judge I had trusted for months.
Software engineer / Miami, FL
The notebook
Small thoughts, things I learned the hard way, and ideas that needed more room.
Follow via RSS — every new thought, field note, and essay in one feed.
Question the measurement
What happened when I tested the judge I had trusted for months.
Follow a fix
An exact failure, the hardware behind it, and the two-line fix.
Look at the craft
The engineering behind a quieter page transition on this site.
3 pieces about Agents & tooling
I have swapped the local model under my newspaper 85 times in four months. That looked like a natural experiment, so I built an instrument to read it — including a gate that refuses comparisons contaminated by time. The gate passed. It was reading my switch log instead of my writing, and the two models it cleared had never once alternated.
I used a language model to pick the better of two drafts, and trusted it for months. Then I gave it two drafts from an identical configuration and asked it to choose. Here is what measuring a judge’s noise floor costs, and why every A/B result before it was unreadable.
Three times this year a reasoning model spent its whole token budget thinking and handed back empty content under HTTP 200. Three times, the code reading the result invented a different explanation — a parser bug, then an unconstrained grammar. Neither was true, and one of them fired inside a safety check.