arXiv:2603.04403cs.IRcs.AI2026-03被引 4

首个评估AI金融数据检索能力的基准,揭示工具使用决定性能上限。

FinRetrieval: A Benchmark for Financial Data Retrieval by AI Agents

  • 构建500个金融问答数据集,包含真实工具调用轨迹与多模型响应。
  • 使用结构化API时,Claude Opus准确率达90.8%,仅用网页搜索时降至19.8%。
  • 工具可用性远超推理模式影响,适合研究金融AI代理与系统设计者参考。

AI代理在金融研究中日益重要,但缺乏评估其从结构化数据库中检索特定数值能力的基准。我们提出FinRetrieval,一个包含500个金融检索问题的数据集,附带真实答案、来自三个领先提供商(Anthropic、OpenAI、Google)14种配置的代理响应,以及完整的工具调用执行轨迹。评估显示,工具可用性主导性能:Claude Opus使用结构化API时准确率达90.8%,仅依赖网页搜索时降至19.8%,差距达71个百分点,超出其他供应商3-4倍。推理模式增益与基础能力呈反比(OpenAI提升9.0个百分点,Claude仅2.8个百分点),源于基础模式工具使用差异而非推理能力。地理性能差距(美国领先5.6个百分点)源于财务年度命名惯例,非模型局限。我们开源数据集、评估代码与工具轨迹,以推动金融AI系统研究。

原文摘要 · Abstract (English)

AI agents increasingly assist with financial research, yet no benchmark evaluates their ability to retrieve specific numeric values from structured databases. We introduce FinRetrieval, a benchmark of 500 financial retrieval questions with ground truth answers, agent responses from 14 configurations across three frontier providers (Anthropic, OpenAI, Google), and complete tool call execution traces. Our evaluation reveals that tool availability dominates performance: Claude Opus achieves 90.8% accuracy with structured data APIs but only 19.8% with web search alone--a 71 percentage point gap that exceeds other providers by 3-4x. We find that reasoning mode benefits vary inversely with base capability (+9.0pp for OpenAI vs +2.8pp for Claude), explained by differences in base-mode tool utilization rather than reasoning ability. Geographic performance gaps (5.6pp US advantage) stem from fiscal year naming conventions, not model limitations. We release the dataset, evaluation code, and tool traces to enable research on financial AI systems.

金融AI检索基准工具调用大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。