
GitHub - bilikaz/qwen38-flash-next-recipe: qwen38-flash-next recipe for solo GB10
qwen38-flash-next recipe for solo GB10. Contribute to bilikaz/qwen38-flash-next-recipe development by creating an account on GitHub.
The repo provides shell scripts to serve the unsloth/Qwen3.8-27B-NVFP4 checkpoint via vLLM in Docker on NVIDIA DGX Spark (GB10, aarch64) or RTX 6000 PRO.
It uses NVFP4 4-bit quantization, FP8 KV cache, YaRN-extended 1M-token context, and MTP speculative decoding. Three scripts handle download, server start (OpenAI-compatible API on port 8888), and stop.