首个开源金融搜索与推理基准,评估大模型真实分析师工作流表现。
FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and Reasoning
- 构建三类真实金融分析任务,模拟多步搜索与时间敏感数据处理。
- 635个问题覆盖全球与大中华市场,Grok 4(web)在全局领先,DouBao(web)在大中华区最优。
- 验证网页搜索与金融插件对性能提升关键,模型来源显著影响表现。
搜索已成为基于大模型代理的核心基础设施,被广泛认为是迈向通用智能的关键路径。金融领域尤其适合作为测试场:分析师需对时效性强、领域专精的数据进行复杂、多步骤的搜索,极好地检验搜索能力与知识驱动推理。然而,现有公开金融数据集大多无法评估端到端代理的数据搜索能力,主要因构建真实、复杂的任务需要深厚的金融专业知识,且时间敏感数据难以评测。我们提出FinSearchComp,首个完全开源的端到端金融搜索与推理代理基准。该基准包含三项任务——时间敏感数据获取、简单历史查询与复杂历史调查——紧密复现真实金融分析师的工作流程。为确保难度与可靠性,我们邀请70位专业金融专家参与标注,并实施多阶段质量保障流程。基准涵盖635个问题,覆盖全球及大中华市场,我们对21个模型(产品)进行了评估。Grok 4 (web) 在全球子集上表现最佳,接近专家水平;DouBao (web) 在大中华子集上领先。实验分析表明,为代理配备网页搜索和金融插件可显著提升性能,且模型与工具的国家来源对表现有显著影响。通过贴合真实分析师任务并提供端到端评估,FinSearchComp为复杂金融搜索与推理提供了专业、高难度的测试平台。
原文摘要 · Abstract (English)
Search has emerged as core infrastructure for LLM-based agents and is widely viewed as critical on the path toward more general intelligence. Finance is a particularly demanding proving ground: analysts routinely conduct complex, multi-step searches over time-sensitive, domain-specific data, making it ideal for assessing both search proficiency and knowledge-grounded reasoning. Yet no existing open financial datasets evaluate data searching capability of end-to-end agents, largely because constructing realistic, complicated tasks requires deep financial expertise and time-sensitive data is hard to evaluate. We present FinSearchComp, the first fully open-source agent benchmark for realistic, open-domain financial search and reasoning. FinSearchComp comprises three tasks -- Time-Sensitive Data Fetching, Simple Historical Lookup, and Complex Historical Investigation -- closely reproduce real-world financial analyst workflows. To ensure difficulty and reliability, we engage 70 professional financial experts for annotation and implement a rigorous multi-stage quality-assurance pipeline. The benchmark includes 635 questions spanning global and Greater China markets, and we evaluate 21 models (products) on it. Grok 4 (web) tops the global subset, approaching expert-level accuracy. DouBao (web) leads on the Greater China subset. Experimental analyses show that equipping agents with web search and financial plugins substantially improves results on FinSearchComp, and the country origin of models and tools impact performance significantly.By aligning with realistic analyst tasks and providing end-to-end evaluation, FinSearchComp offers a professional, high-difficulty testbed for complex financial search and reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。