Welcome! Type "help" for available commands.
$
Pre-built Docker/Toolbx/Distrobox container for running the ds4 DeepSeek V4 Flash inference engine on AMD Strix Halo integrated GPUs (gfx1151) with ROCm 7.14.
Exposes ds4, ds4-server, and ds4-bench binaries for interactive chat and OpenAI/Anthropic-compatible HTTP serving. Runs imatrix-quantized GGUF models (IQ2_XXS, hybrid Q2/Q4) using up to 124 GiB unified memory at contexts up to 124,000 tokens.