Mellum2-12B-A2.5B-Thinking is a JetBrains reasoning-augmented code assistant utilizing a Mixture-of-Experts architecture with 12B total parameters and a 131K-token context. It explicitly emits chain-of-thought reasoning traces within tags, making it ideal for debugging and multi-step planning.
Highlights
Supports a 131,072-token context window using a hybrid of sliding-window and full-attention layers.
Requires vLLM nightly builds (post-v0.22.0) for MellumForCausalLM support and specific reasoning parsers like qwen3.
JetBrains recommends the Instruct variant for low-latency direct answers, reserving the Thinking variant for complex reasoning tasks.
Jackrong/Qwopus3.6-27B-Coder-MTP-GGUF · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
github.com
GitHub - centminmod/or-cli: Python command-line tool for interacting with AI models through the OpenRouter API/Cloudflare AI Gateway, or local self-hosted Ollama. Optionally support Microsoft LLMLingua prompt token compression
Python command-line tool for interacting with AI models through the OpenRouter API/Cloudflare AI Gateway, or local self-hosted Ollama. Optionally supp...
huggingface.co
Qwen/Qwen3-235B-A22B-Instruct-2507 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.