Why Prefill and Decode Should Not Share a GPU
A systems-oriented introduction to Prefill-Decode disaggregation in LLM inference, from KV cache physics and roofline intuition to vLLM and SGLang implementation trade-offs.
Content tagged with "inference"
A systems-oriented introduction to Prefill-Decode disaggregation in LLM inference, from KV cache physics and roofline intuition to vLLM and SGLang implementation trade-offs.