|
Mila
Deep Neural Network Library
|
KV cache compression policy concept and identity struct. More...
#include <concepts>Classes | |
| struct | Mila::Dnn::Quant::KvCache::NoKvCompression |
| Identity policy - no compression. More... | |
| struct | Mila::Dnn::Quant::KvCache::SlidingWindowKvCache |
| Bounded sliding-window KV cache (uncompressed ring buffer). More... | |
Namespaces | |
| namespace | Mila |
| Mila main API namespace. | |
Concepts | |
| concept | Mila::Dnn::Quant::KvCache::KvCachePolicy |
| Minimum contract for all KV cache management policies. | |
KV cache compression policy concept and identity struct.
Establishes the KvCachePolicy extension point consumed by GroupedQueryAttention and CudaGqaOp. All active compression policies (PerChannelKvFp8, future SlidingWindow, future MLA) satisfy this concept. GroupedQueryAttention constrains its TKvPolicy parameter to KvCachePolicy.
The concept is intentionally minimal: only kIsActive is required. Active policy refinements (QuantKvPolicy) impose additional field requirements in their own modules. A SlidingWindowPolicy satisfies KvCachePolicy without carrying dtype fields at all.