
RedHatAI/gemma-4-31B-it-NVFP4 · Hugging Face
NVFP4 variant of gemma-4-31B-it.
QAT-quantized GGUF formats (Q4_0, Q4_K_XL) of Google's Gemma 4 26B A4B MoE model (25.2B total, 3.8B active params) by unsloth.
Ships with a Multi-Token Prediction drafter for llama.cpp speculative decoding that shares the target KV cache. Supports text and image input, 256K context, Apache 2.0 license.