Articles
Prefill and decode in LLM inference
An overview of prefill and decode, the two phases of ordinary LLM inference, and why one tends to be bottlenecked by compute while the other tends to be bottlenecked by memory reads — prompted by sessions at KubeCon + CloudNativeCon Japan 2026.