Welcome! Type "help" for available commands.
$
vLLM serving recipe for GLM-5.3-Flash using EXL3/TR3 4-bpw quantization (~164 GiB) on 2x NVIDIA GB10 with tensor-parallel size 2 over CX7.
Serves an OpenAI-compatible API on port 8888 with 1M context, fp8 packed KV cache, and DFlash2 k=7 speculative decoding. Weights are a byte-identical Hugging Face mirror of brandonmusic's EXL3 checkpoint to keep the recipe fetchable independently of the upstream Hub ID.