FinFIRST is a specialized financial benchmark designed to evaluate LLM search agents on both their final answers and the traceability of their supporting evidence. It utilizes atomic rubrics to assess information acquisition, source verification, and computation across 123 expert-authored tasks.
Highlights
Evaluates both final answers and supporting evidence through fine-grained atomic rubrics
Comprises 123 tasks built from an 18-field taxonomy and 138 registered financial sources
Identifies computation and answer formation as a consistent performance bottleneck for LLM agents
Provides a measurable and diagnosable framework for the entire financial research process
LLM AgentsFinancial Information RetrievalAI BenchmarkingEvidence Grounding
Discover Similar Content
huggingface.co
inclusionAI/FinFIRST · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
vals.ai
Vals AI
Private, domain-specific benchmarks in legal, tax, and finance.
github.com
GitHub - centminmod/or-cli: Python command-line tool for interacting with AI models through the OpenRouter API/Cloudflare AI Gateway, or local self-hosted Ollama. Optionally support Microsoft LLMLingua prompt token compression
Python command-line tool for interacting with AI models through the OpenRouter API/Cloudflare AI Gateway, or local self-hosted Ollama. Optionally supp...