Andrea Pellegrini details building a local LLM inference stack on the AMD Ryzen AI Max+ 395 with 128GB unified memory. The article covers performance benchmarks using Vulkan and ROCm backends for models up to 142B parameters.
Highlights
Utilization of AMD Ryzen AI Max+ 395 with 128GB unified memory for local LLM inference.
Performance benchmarks showing up to 884 tokens/s for Llama 2 7B Q4_0 using Vulkan.
Achievement of 270 tokens/s for 120B parameter models at pp512 context length.
Comparison of different inference backends including HIP, Vulkan, and ROCm.
Demonstration of running large models up to 142B weights through hardware optimization.
Backend benchmarks for AMD Strix Halo measuring tokens-per-second throughput across 17 LLM configurations ranging from 6.7B to 228.7B parameters. Test...
youtube.com
GLM 4.5-Air-106B and Qwen3-235B on AMD “Strix Halo” AI Ryzen MAX+ 395 (HP Z2 G1a Mini Workstation)
In this video, I show how to run large language models like Qwen3-235B in Q3_K_XL and GLM 4.5-Air-106B in Q6_K_XL on any AMD Ryzen AI MAX “Strix Halo”...
github.com
GitHub - julianmb/q38rocm: Qwen 3.8 27B ROCmFP4 on AMD Strix Halo (Ryzen AI Max+ 395). Up to 36 tok/s via MTP Speculation, TurboQuant & Mesa RADV Wave64.
Qwen 3.8 27B ROCmFP4 on AMD Strix Halo (Ryzen AI Max+ 395). Up to 36 tok/s via MTP Speculation, TurboQuant & Mesa RADV Wave64. - julianmb/q38rocm