Text Generation
Safetensors
Model Optimizer
gemma4
nvidia
ModelOpt
Gemma-4-31B-IT
lighthouse
quantized
NVFP4
conversational
modelopt
Instructions to use nvidia/Gemma-4-31B-IT-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Inference
Why not quantize the MATRICES of Wq, Wk, Wv, Wo?
#5
by BeetSoup - opened
The weights \model.language_model.layers.\d+.self_attn.[qkvo]_proj.weight\ make up 28.08% of the model's parameters. However, by keeping them in bf16, u make the model achieves about 8bpw in avg.
I think because quantizing them would significantly degrade the model performance.