This content details an experiment using AI agents for training optimization and model compression, showing that tight scoping prevents task drift. The authors open-sourced their bash-wrapped Codex loop and techniques for compressing a 2.5 TB model to fit consumer GPUs.
Highlights
Autoresearch agents require tight scoping to avoid drifting into unrelated tasks.
The authors open-sourced a bash-wrapped Codex loop enforcing A/B testing.
Experiment 2 compressed a 2.5 TB Kimi-k2.5 model to fit on 8x RTX 3090s.
The bottleneck for autonomous AI research is infrastructure, not intelligence.
Testing different LLMs revealed varying behaviors in autonomous research loops.
GitHub - karpathy/autoresearch: AI agents running research on single-GPU nanochat training automatically
AI agents running research on single-GPU nanochat training automatically - karpathy/autoresearch
github.com
GitHub - uditgoenka/autoresearch: Claude Autoresearch Skill — Autonomous goal-directed iteration for Claude Code. Inspired by Karpathy’s autoresearch. Modify → Verify → Keep/Discard → Repeat forever.
Claude Autoresearch Skill — Autonomous goal-directed iteration for Claude Code. Inspired by Karpathy’s autoresearch. Modify → Verify → Keep/Discard → ...
github.com
GitHub - TheGreenCedar/codex-autoresearch: A codex plugin for running optimization loops inside a codebase. It is useful when you have a measurable target and many possible changes to try: test runtime, build speed, bundle size, model loss, Lighthouse scores, memory use, query latency, or any other metric you can print from a script.
A codex plugin for running optimization loops inside a codebase. It is useful when you have a measurable target and many possible changes to try: test...