writing
Technical articles on quantization, speculative decoding, and LLM serving.
Technical notes connecting mathematical derivations to working implementations. Additional writing is on Medium.
June 2026
Unpacking Speculative Decoding: The Math Behind the Speedup
A derivation of speculative decoding acceptance probability, expected accepted tokens, and the speedup objective.
Read articleJune 2026
The Math of AWQ: Protecting Salient Channels from the Inside Out
A first principles explanation of activation aware weight quantization and its scaling tradeoffs.
Read articleJune 2026
Demystifying GPTQ: From Lagrange Multipliers to Vectorized PyTorch
A derivation of the inverse Hessian update used by GPTQ, followed by a vectorized PyTorch implementation.
Read article