Factory Research evaluated context compression for AI agents using probe-based tests. They found structured summarization preserves more details than OpenAI or Anthropic methods during long sessions.
Highlights
Traditional metrics like ROUGE fail to measure functional context preservation.
Probe-based tests verify if agents recall specific details after compression.
Structured summarization outperformed OpenAI and Anthropic in debugging tasks.
Optimizing for tokens per task improves agent productivity over tokens per request.
Testing covered debugging, code review, and ML research scenarios.
Context Compaction: Query-Based and Objective Modes | Morph
Compaction drops 50-70% of an agent’s context while keeping every surviving line verbatim. Two modes: objective compaction strips filler with no guida...
github.com
GitHub - centminmod/or-cli: Python command-line tool for interacting with AI models through the OpenRouter API/Cloudflare AI Gateway, or local self-hosted Ollama. Optionally support Microsoft LLMLingua prompt token compression
Python command-line tool for interacting with AI models through the OpenRouter API/Cloudflare AI Gateway, or local self-hosted Ollama. Optionally supp...
warp.dev
Skill Doctor: score and improve your agent skills
Score past Claude Code, Codex, and Warp agent conversations against tested rubrics, then get concrete edits to your SKILL.md files — free and open sou...