Mila
Deep Neural Network Library
Loading...
Searching...
No Matches
Dnn.LanguageModelConfig Module Reference

Classes

struct  Mila::Dnn::LanguageModelConfig< TDerived >
 CRTP base configuration for all deployable Mila language models. More...

Enumerations

enum class  Mila::Dnn::KvCacheCompression { None , FP8 }
 KV cache storage and compression strategy for GroupedQueryAttention. More...
enum class  Mila::Dnn::WeightQuantization { None , FP8 , FP4 }
 Weight storage and matmul strategy for Linear components. More...

Functions

std::string Mila::Dnn::weightQuantizationName (WeightQuantization quantization)
 The scheme name recorded in an artifact and in its manifest.

Files

file  Mila/Src/Dnn/Core/LanguageModelConfig.ixx
 CRTP base configuration for all deployable Mila language models.