|
Mila
Deep Neural Network Library
|
Classes | |
| struct | Mila::Dnn::LanguageModelConfig< TDerived > |
| CRTP base configuration for all deployable Mila language models. More... | |
Enumerations | |
| enum class | Mila::Dnn::KvCacheCompression { None , FP8 } |
| KV cache storage and compression strategy for GroupedQueryAttention. More... | |
| enum class | Mila::Dnn::WeightQuantization { None , FP8 , FP4 } |
| Weight storage and matmul strategy for Linear components. More... | |
Functions | |
| std::string | Mila::Dnn::weightQuantizationName (WeightQuantization quantization) |
| The scheme name recorded in an artifact and in its manifest. | |
Files | |
| file | Mila/Src/Dnn/Core/LanguageModelConfig.ixx |
| CRTP base configuration for all deployable Mila language models. | |