|
Mila
Deep Neural Network Library
|
Polymorphic inference interface for one decoder layer. More...
Public Types | |
| using | MR = typename DeviceTypeTraits<TDeviceType>::memory_resource |
| using | TensorType = Tensor<TPrecision, MR> |
Public Member Functions | |
| virtual TensorType & | decode (const TensorType &input, dim_t position)=0 |
| Single-token decode at an absolute position (T == 1). | |
| virtual TensorType & | prefill (const TensorType &input, dim_t position_offset)=0 |
| Chunked prefill: process [B, T_chunk, model_dim] at an absolute offset. | |
| virtual void | resetKVCache ()=0 |
| Reset the KV cache (new generation session). | |
| virtual bool | rewindKvCache (dim_t position)=0 |
| Rewind the KV cache fill position for prompt-prefix reuse. | |
| virtual void | setState (const GqaState &state)=0 |
| Wire the shared GQA transient workspace (owned by the transformer). | |
| virtual bool | supportsKVCache () const noexcept=0 |
| True when the block's attention supports the KV-cache inference path. | |
Polymorphic inference interface for one decoder layer.
| TDeviceType | Compile-time device. |
| TPrecision | Activation/compute precision (must match across the layer list). |
|
pure virtual |
Single-token decode at an absolute position (T == 1).
Implemented in Mila::Dnn::GemmaBlock< TDeviceType, TPrecision, kGlobal, TWeightQuant, TKvPolicy >, Mila::Dnn::GemmaBlock< TDeviceType, TPrecision, false, TWeightQuantization, TKvCachePolicy >, and Mila::Dnn::GemmaBlock< TDeviceType, TPrecision, true, TWeightQuantization, NoKvCompression >.
|
pure virtual |
Chunked prefill: process [B, T_chunk, model_dim] at an absolute offset.
Implemented in Mila::Dnn::GemmaBlock< TDeviceType, TPrecision, kGlobal, TWeightQuant, TKvPolicy >, Mila::Dnn::GemmaBlock< TDeviceType, TPrecision, false, TWeightQuantization, TKvCachePolicy >, and Mila::Dnn::GemmaBlock< TDeviceType, TPrecision, true, TWeightQuantization, NoKvCompression >.
|
pure virtual |
Reset the KV cache (new generation session).
Implemented in Mila::Dnn::GemmaBlock< TDeviceType, TPrecision, kGlobal, TWeightQuant, TKvPolicy >, Mila::Dnn::GemmaBlock< TDeviceType, TPrecision, false, TWeightQuantization, TKvCachePolicy >, and Mila::Dnn::GemmaBlock< TDeviceType, TPrecision, true, TWeightQuantization, NoKvCompression >.
|
pure virtual |
Rewind the KV cache fill position for prompt-prefix reuse.
Keeps the cache session live; positions [0, position) stay valid.
Implemented in Mila::Dnn::GemmaBlock< TDeviceType, TPrecision, kGlobal, TWeightQuant, TKvPolicy >, Mila::Dnn::GemmaBlock< TDeviceType, TPrecision, false, TWeightQuantization, TKvCachePolicy >, and Mila::Dnn::GemmaBlock< TDeviceType, TPrecision, true, TWeightQuantization, NoKvCompression >.
|
pure virtual |
Wire the shared GQA transient workspace (owned by the transformer).
Called once after build, before any prefill/decode.
Implemented in Mila::Dnn::GemmaBlock< TDeviceType, TPrecision, kGlobal, TWeightQuant, TKvPolicy >, Mila::Dnn::GemmaBlock< TDeviceType, TPrecision, false, TWeightQuantization, TKvCachePolicy >, and Mila::Dnn::GemmaBlock< TDeviceType, TPrecision, true, TWeightQuantization, NoKvCompression >.
|
pure virtualnoexcept |
True when the block's attention supports the KV-cache inference path.
Implemented in Mila::Dnn::GemmaBlock< TDeviceType, TPrecision, kGlobal, TWeightQuant, TKvPolicy >, Mila::Dnn::GemmaBlock< TDeviceType, TPrecision, false, TWeightQuantization, TKvCachePolicy >, and Mila::Dnn::GemmaBlock< TDeviceType, TPrecision, true, TWeightQuantization, NoKvCompression >.