writing
Technical articles on quantization, speculative decoding, and LLM serving.
Technical notes connecting mathematical derivations to working implementations. Additional writing is on Medium.
August 2026
Medusa Part II: Candidate Trees, Verification, and Acceptance
A complete runtime guide to Medusa candidate pools, sparse rank trees, tree attention, greedy and typical acceptance, and an A100 trace.
Read articleAugust 2026
Medusa Part I: Architecture, Training, and Shifted Loss
A tensor-level guide to Medusa heads, frozen and joint training, label shifts, and the complete shifted cross-entropy calculation.
Read articleJune 2026
Unpacking Speculative Decoding: The Math Behind the Speedup
A derivation of speculative decoding acceptance probability, expected accepted tokens, and the speedup objective.
Read articleJune 2026
The Math of AWQ: Protecting Salient Channels from the Inside Out
A first principles explanation of activation aware weight quantization and its scaling tradeoffs.
Read articleJune 2026
Demystifying GPTQ: From Lagrange Multipliers to Vectorized PyTorch
A derivation of the inverse Hessian update used by GPTQ, followed by a vectorized PyTorch implementation.
Read article