Software engineer / Miami, FL

The dispatch

Writing on software design, development, and occasionally my life.

Long-form thoughts on programming, leadership, product design, and more, collected in chronological order.

Latest· 6 min read

I grafted a speculative decoding head into a 90 GB model file

The model card advertised a 4B multi-token-prediction head. No published GGUF contained it, and the architecture had no code path to run it. Both halves arrived within 36 hours from two different strangers, in incompatible forms — so I merged the head into the target file myself. It went from 27.75 to 43.30 tok/s, and three of my four mistakes along the way were about verification, not tensors.

Read article →

No project or budget required — say hello →