This repository provides optimized deployment for the Qwen 3.8 127B model on AMD Strix Halo APUs via custom ROCmFP4 quantization. Utilizing MTP speculative decoding and RADV Wave62 engines, it reaches speeds of 36 tokens per second by overcoming memory bandwidth bottlenecks.
Highlights
Achieves high throughput of up to 36 tokens per second
Uses custom ROCmFP4 block quantization designed for RDNA 3.5
Requires a specific ROCmFPX-enabled version of llama.cpp
Optimized specifically for AMD Strix Halo unified memory architecture
auto-generated
julianmb Β· via GitHub
Context
Audience
Machine Learning Engineers and AI Hardware Enthusiasts