Inference
Understanding FP8 KV Cache in vLLM
Understand what KV cache quantization changes and build a compatibility and quality check before enabling it.
Documentation-based guideTechnical notes organized around a shared topic.
Understand what KV cache quantization changes and build a compatibility and quality check before enabling it.
Documentation-based guideDistinguish the attention state stored during generation from reuse of that state across requests.
Documentation-based guide