Welcome! Type "help" for available commands.
$
An MLX-based inference engine for Apple Silicon that evaluates multi-field JSON schemas in parallel via KV-cache broadcasting, achieving 5.6x to 7.0x latency reduction over autoregressive decoding with 100% schema validity.
Supports structured extraction, decision routing, and categorical classification with up to 255 enum choices per field and calibrated confidence scores. Includes a Python SDK, M4 Max performance benchmarks, and a live web visualizer for side-by-side comparison.