Welcome! Type "help" for available commands.
$
Off-policy and on-policy distillation recipes for post-training with Tinker, built on OpenThoughts3, DeepMath, and Tulu3 datasets.
Achieves ~65% AIME'24 via supervised fine-tuning (rank-128 LoRA, 3000 steps) and ~76.7% via on-policy distillation (rank-128 LoRA, 200 steps, 16k-token rollouts). Includes multi-turn tool-use distillation using Harbor sandbox environments with KL divergence against a teacher as the sole training signal.