arXiv:2508.14052cs.IRcs.AI2025-08被引 35

首个面向金融领域智能检索的基准数据集,评测大模型多步推理能力。

FinAgentBench: A Benchmark Dataset for Agentic Retrieval in Financial Question Answering

  • 构建金融领域多步推理检索框架,分离文档类型与关键段落识别任务
  • 包含26,000个专家标注样本,覆盖标普500上市公司信息
  • 适用于评估大模型在复杂金融场景下的检索行为,支持模型优化

精准的信息检索(IR)在金融领域至关重要,投资者需从海量文档中识别相关资讯。传统稀疏或稠密检索方法因难以兼顾语义相似性与文档结构、领域知识的细粒度推理,常表现不足。大语言模型(LLMs)的兴起为多步推理检索带来新可能,模型可通过迭代推理对段落进行排序。然而,金融领域尚无此类能力的评测基准。为此,我们提出FinAgentBench,首个针对金融领域多步推理检索的大型基准数据集——我们称之为“智能体式检索”。该数据集包含26,000个专家标注样例,涵盖标普500上市公司信息,评估LLM智能体能否(1)从候选文档中识别最相关类型,(2)精确定位所选文档中的关键段落。评估框架明确分离这两个推理步骤,以克服上下文限制。该设计为理解金融领域以检索为中心的LLM行为提供了量化基础。我们测试了多种前沿模型,并证明针对性微调可显著提升智能体式检索性能。本基准为研究复杂领域任务中以检索为核心的LLM行为奠定了基础。

原文摘要 · Abstract (English)

Accurate information retrieval (IR) is critical in the financial domain, where investors must identify relevant information from large collections of documents. Traditional IR methods -- whether sparse or dense -- often fall short in retrieval accuracy, as it requires not only capturing semantic similarity but also performing fine-grained reasoning over document structure and domain-specific knowledge. Recent advances in large language models (LLMs) have opened up new opportunities for retrieval with multi-step reasoning, where the model ranks passages through iterative reasoning about which information is most relevant to a given query. However, there exists no benchmark to evaluate such capabilities in the financial domain. To address this gap, we introduce FinAgentBench, the first large-scale benchmark for evaluating retrieval with multi-step reasoning in finance -- a setting we term agentic retrieval. The benchmark consists of 26K expert-annotated examples on S&P-500 listed firms and assesses whether LLM agents can (1) identify the most relevant document type among candidates, and (2) pinpoint the key passage within the selected document. Our evaluation framework explicitly separates these two reasoning steps to address context limitations. This design enables to provide a quantitative basis for understanding retrieval-centric LLM behavior in finance. We evaluate a suite of state-of-the-art models and further demonstrated how targeted fine-tuning can significantly improve agentic retrieval performance. Our benchmark provides a foundation for studying retrieval-centric LLM behavior in complex, domain-specific tasks for finance.

金融AI检索增强大模型评测多步推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。