
GitHub - MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark: Qwen3.8-Flash-Next on ONE DGX Spark (TP=1)
Qwen3.8-Flash-Next on ONE DGX Spark (TP=1). Contribute to MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark development by creating an account on GitHub.
Alok benchmarks Gemma 4 26B A4B MoE on a single RTX 4090 (24 GB) using Llama.cpp 8-bit KV cache quantization, sustaining 24 concurrent users at 500 t/s aggregate decode within 23.35 GB VRAM.
Unquantized KV cache for the same 24-user load exceeds 28 GB and triggers OOM. The tradeoff is prefill speed dropping from roughly 1,500 t/s to 750 t/s, exchanging latency for a 71% capacity increase.