Welcome! Type "help" for available commands.
$
Custom ROCmFP4 and ROCmFP8 quantized GGUF builds of the Muse-Glimmer-30B multimodal model, tested on AMD Strix Halo (gfx1151).
Files include main-model quantizations (14.17–26.77 GiB), a BF16 vision projector, and DFlash drafters; they require a patched ROCmFPX runtime and will not load in stock llama.cpp. Measured prompt throughput ranges from 39 to 113.7 tokens/s with output rates of 7.8 to 28.3 tokens/s depending on quantization and drafter configuration.