Welcome! Type "help" for available commands.
$
FINFIRST is a 123-task financial benchmark that scores LLM search agents on both final answers and supporting evidence via atomic rubrics covering information acquisition, source verification, and computation.
Tasks are built from an 18-field taxonomy, 138 registered financial sources, and a 6-stage quality control pipeline. Across 15 model configurations, Claude-Opus-5 achieved the top atomic score (87.59%) while GPT-5.6-Sol led in strict pass rate (71.54%), with computation lagging as a common bottleneck.