
unsloth/Kimi-K2-Instruct-GGUF · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Trinity Mini is a 26B-parameter sparse mixture-of-experts language model with 3B active parameters, designed for efficient reasoning and multi-step agent workflows over extended contexts up to 131k tokens.
The model features 128 total experts with 8 active per token and supports robust function calling capabilities. Trinity Mini is available free through OpenRouter with 100% uptime, achieving approximately 213 tokens per second throughput and 0.32 seconds latency.