
Overview - Together AI docs
Adapt a base model to a task by training it on your data.
A GitHub repository benchmarking LoRA fine-tuning speed: reach ≥57% GSM8K accuracy on Qwen2.5-1.5B on a single NVIDIA L40S, tracked on a public wall-clock leaderboard.
Submissions use modded-nanogpt with frozen base weights and a ≤30M-parameter PEFT adapter; records require three fresh-seed verification runs in a network-blocked Modal sandbox. The current record is 6 minutes 05 seconds via sequence packing and completion-only loss masking over 2 epochs.