FreeToken is an edge-native Mixture-of-Experts (MoE) serving engine designed to run massive, frontier-scale open-weight models on consumer-grade hardware like gaming PCs and laptops. It utilizes bandwidth-adaptive execution and elastic memory management to provide datacenter-class intelligence on personal devices.
Highlights
Enables running 290B+ parameter MoE models locally on consumer GPUs (NVIDIA RTX 30, 40, and 50 series).
Features semantic-aware caching and elastic memory management to optimize VRAM usage without engine restarts.
Supports a wide range of frontier models including DeepSeek-V4-Flash, Qwen, and GLM across various quantization formats.
Provides both a user-friendly desktop GUI and a command-line interface (CLI) for flexible deployment.
Uses a bandwidth-adaptive CPU–GPU co-execution policy for efficient edge-native runtime performance.
auto-generated
FlashML-org · via GitHub
Context
Audience
AI Researchers, Machine Learning Engineers, and Developers with consumer-grade hardware