LLM
KV Cache Offloading Strategies and a Practical Understanding of GQA
Hello! Welcome to a special extra installment of our LLM Inference Infrastructure Provisioning Course! Part 3, "Estimating Inference-Time Memory Consumption for Your Model", we introduced the two major consumers of GPU memory, the model footprint and the KV cache, and explained how to calculate the KV cache size