
GitHub - bilikaz/qwen38-flash-next-recipe: qwen38-flash-next recipe for solo GB10
qwen38-flash-next recipe for solo GB10. Contribute to bilikaz/qwen38-flash-next-recipe development by creating an account on GitHub.
vLLM serving recipe for GLM-5.3-Flash using EXL3/TR3 4-bpw quantization (~164 GiB) on 2x NVIDIA GB10 with tensor-parallel size 2 over CX7.
Serves an OpenAI-compatible API on port 8888 with 1M context, fp8 packed KV cache, and DFlash2 k=7 speculative decoding. Weights are a byte-identical Hugging Face mirror of brandonmusic's EXL3 checkpoint to keep the recipe fetchable independently of the upstream Hub ID.