Welcome! Type "help" for available commands.
$
NVFP4 quantized variant of google/gemma-4-31B-it using FP4 for weights (group_size=16) and activations, reducing memory by ~75%.
Vision tower, embeddings, and output head retain original precision; deployable via vLLM with multimodal text/image input. Benchmarked on GSM8K, MMLU-Pro, IFEval, MATH-500, AIME 2025, GPQA Diamond, LiveCodeBench, and BFCLv4 with 95–100% accuracy recovery.