Welcome! Type "help" for available commands.
$
Serves a 99 GiB Qwen3.8-Flash-Next-NVFP4 vision-language model (text, image, video) from a single DGX Spark's 121 GiB unified memory via vLLM with the PLE table memory-mapped.
Ships with 262k context, FP8 KV cache, MTP 3, and up to 8 concurrent sequences; measured at 48.7 tok/s single-stream and 162.9 tok/s aggregate at 8 streams. Requires ~130 GiB free disk, reaches health in 10-12 minutes, and serves an API on port 8888.