A highly efficient 2-bit quantized version of the Qwen3.8-27B large language model designed to run on consumer-grade GPUs. It enables a 27B parameter model to fit within approximately 10.15 GB of VRAM, allowing for large context windows on hardware like the RTX 3090 or 4090.
Highlights
Uses mixed 2/3-bit quantization to achieve a small 10.15 GB weight footprint
Optimized to fit a 27B model and 64k context on a single 24 GB consumer GPU
Maintains competitive performance in commonsense reasoning and coding compared to FP8 references
Requires a specialized SGLang-based runtime (escha-runtime-qwen3dense) for deployment
Qwen3.8-27BSGLangQuantizationNVIDIA CUDAPyTorchHugging Face
Discover Similar Content
docs.sglang.io
Qwen3.8-27B - SGLang Documentation
Deploy Qwen3.8-27B with SGLang — dense hybrid GDN vision-language model with BF16/FP8/NVFP4 W4A4 checkpoints and in-checkpoint MTP, single-GPU on H200...
github.com
GitHub - julianmb/q38rocm: Qwen 3.8 27B ROCmFP4 on AMD Strix Halo (Ryzen AI Max+ 395). Up to 36 tok/s via MTP Speculation, TurboQuant & Mesa RADV Wave64.
Qwen 3.8 27B ROCmFP4 on AMD Strix Halo (Ryzen AI Max+ 395). Up to 36 tok/s via MTP Speculation, TurboQuant & Mesa RADV Wave64. - julianmb/q38rocm
huggingface.co
outsourc-e/Qwen3.8-27B-Unleashed-GGUF · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.