Inference
Understanding FP8 KV Cache in vLLM
Understand what KV cache quantization changes and build a compatibility and quality check before enabling it.
Documentation-based guideDeploy, serve and optimize. Practical notes on the engines behind LLM inference.
Understand what KV cache quantization changes and build a compatibility and quality check before enabling it.
Documentation-based guideDistinguish the attention state stored during generation from reuse of that state across requests.
Documentation-based guideChoose an inference engine using workload compatibility, reproducibility and operational requirements—not an unqualified leaderboard.
Documentation-based guide