a 9-part series on testing LLM products, from why-evals to multi-turn agents.
an inference-engineering primer. what quantization is, why it works, which formats matter in 2026 (fp8, fp4, mxfp8, nvfp4, nf4), the five mainstream techniques with working code, kv-cache quantization, the sensitivity hierarchy, and a safe production recipe for h100 / b200.
an inference-engineering primer. what quantization is, why it works, which formats matter in 2026 (fp8, fp4, mxfp8, nvfp4, nf4), the five mainstream techniques with working code, kv-cache quantization, the sensitivity hierarchy, and a safe production recipe for h100 / b200.
why setTimeout(0) is never zero, why await feels seamless, and why one runaway Promise can stall a tab.
why setTimeout(0) is never zero, why await feels seamless, and why one runaway Promise can stall a tab.