
outsourc-e/Qwen3.8-27B-Unleashed-GGUF · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
GGUF quantizations of Qwen3.8-Flash-Next, a 512-expert sparse MoE model, at four bit-widths using per-tensor non-uniform quantization assigned via GSQ and RCO gradient-based search.
Files range from 66.4 GB (Q2_0, 2.40 bpw) to 83.6 GB (IQ3_S, 3.50 bpw), each split into a transformer shard and a 28.8 GB n-gram lookup shard; all run unmodified in llama.cpp, Ollama, and LM Studio. IQ3_S matches or exceeds the base model on every evaluated task; Q2_0 trades 0.09 task-average points for 3.4x prompt throughput and 1.9x lower end-to-end latency.