
JetBrains/Mellum2-12B-A2.5B-Thinking — 12B / 2.5B active · MOE · 128K ctx
JetBrains’ reasoning-augmented code MoE (12B total / 2.5B active) that emits explicit <think> chains for debugging, planning, and agentic coding
GLM-5 is a 744B-parameter MoE model (40B active) from Zhipu AI, scaled up from GLM-4.5's 355B with 28.5T pre-training tokens and DeepSeek Sparse Attention for long-context efficiency.
It outperforms GLM-4.7 on reasoning, coding, and agentic benchmarks like Vending Bench 2 ($4,432 balance) and leads open-source models, nearing Claude Opus 4.5. The model is open-sourced under MIT on Hugging Face and ModelScope, supports local deployment, and integrates with Z.ai for document generation and agent modes.