LaurentZuijdwijk · via GitHub
llama.cpp with adaptive speculative decoding (--spec-draft-adaptive) and a Vulkan backend tuned for AMD Strix Halo. 4.7x on structured output, 1.9x mainline prefill on MoE. - LaurentZuijdwijk/llama…