Welcome! Type "help" for available commands.
$
Alok benchmarks Gemma 4 26B A4B MoE on a single RTX 4090 (24 GB) using Llama.cpp 8-bit KV cache quantization, sustaining 24 concurrent users at 500 t/s aggregate decode within 23.35 GB VRAM.
Unquantized KV cache for the same 24-user load exceeds 28 GB and triggers OOM. The tradeoff is prefill speed dropping from roughly 1,500 t/s to 750 t/s, exchanging latency for a 71% capacity increase.