This paper introduces HANDBOOK.md, a benchmark evaluating how language-model agents adhere to long policy documents during tool-use tasks. It uses 65 enterprise-modeled tasks to test if agents can maintain strict compliance with instructions spanning 20 to 124 pages.
Highlights
Introduces HANDBOOK.md, a benchmark of 65 tasks across five professional domains.
Tests agent adherence to long-form standard operating procedures (20-124 pages).
Uses a deterministic grading system with 824 programmatic criteria.
Identifies failure patterns where environmental requests override standing policies.
Shows that top frontier models struggle, with the best passing only 36.2% of trials.
auto-generated
Liudas Panavas · via arXiv.org
Context
Audience
AI Researchers, Machine Learning Engineers, and LLM Developers