Welcome! Type "help" for available commands.
$
The repo provides shell scripts to serve the unsloth/Qwen3.8-27B-NVFP4 checkpoint via vLLM in Docker on NVIDIA DGX Spark (GB10, aarch64) or RTX 6000 PRO.
It uses NVFP4 4-bit quantization, FP8 KV cache, YaRN-extended 1M-token context, and MTP speculative decoding. Three scripts handle download, server start (OpenAI-compatible API on port 8888), and stop.