Software engineer / Miami, FL

The notebook

Writing

Small thoughts, things I learned the hard way, and ideas that needed more room.

Follow via RSS — every new thought, field note, and essay in one feed.

A few places to start

Browse the notebook

Clear filters

4 pieces about Local models

Essay

The cache hit that answered the wrong question

I turned prefix caching back on for a 125B hybrid model on a DGX Spark, with the fixes carried from five open pull requests. Hits landed, a number planted deep in the prompt came back exactly every time, and an 11,000-token prompt dropped from 5.99 s to 0.96 s. Then one request answered the question two other requests were asking. Proving that was not a leak took a better instrument than the one I started with.

Essay

I grafted a speculative decoding head into a 90 GB model file

The model card advertised a 4B multi-token-prediction head. No published GGUF contained it, and the architecture had no code path to run it. Both halves arrived within 36 hours from two different strangers, in incompatible forms — so I merged the head into the target file myself. It went from 27.75 to 43.30 tok/s, and three of my four mistakes along the way were about verification, not tensors.

Essay

I turned off concurrency and the server got faster

Four slots served my agentic pipeline at 1.2% prefix-cache reuse. One slot served the same pipeline at 86.7%, and saved 11.6 minutes of prefill in a 39-minute window. Slots schedule requests; they do not add compute — and on a single GPU they scatter the one thing a prefill-bound workload actually depends on.

Essay

My LLM judge was flipping a coin a third of the time

I used a language model to pick the better of two drafts, and trusted it for months. Then I gave it two drafts from an identical configuration and asked it to choose. Here is what measuring a judge’s noise floor costs, and why every A/B result before it was unreadable.