Welcome! Type "help" for available commands.
$
A Podman/Docker deployment that serves Qwen3.8-Flash-Next via an OpenAI-compatible API, using kernels written exclusively for AMD Strix Halo (gfx1151).
At 5.53 bits per weight, it completes a 32K-prompt, 256-token generation in 29.1 seconds, roughly 4x faster than published alternatives on the same hardware. Speculative decoding uses both the model's draft head and prompt lookup, producing byte-identical output to serial greedy decode at temperature 0, and the engine ingests standard GGUF files without a conversion step.