
GitHub - warpfront/hipfire: RDNA-native LLM inference engine in Rust.
RDNA-native LLM inference engine in Rust. . Contribute to warpfront/hipfire development by creating an account on GitHub.
This repository provides quantized Qwen3.8-27B model files (MQ3V2 through MQ6V2, plus MQ4L Lloyd-Max tiers) for use with hipfire, a Rust-native LLM inference engine targeting AMD GPUs.
Each bit-width ships three tiers (xt, base, pro) with DFlash drafters for speculative decoding; MQ4V2 is the canonical MQ4 product. Runtime support covers RDNA4 prefill and decode plus RDNA3 decode, with defaults of Q8 KV cache, 262K-token context, and 81K-token max output.