arXiv:2602.07294cs.CEcs.AI2026-02KDD被引 8

构建真实金融分析评估基准,测试大模型跨文档跨时间的财报理解能力。

Fin-RATE: A Real-world Financial Analytics and Tracking Evaluation Benchmark for LLMs on SEC Filings

  • 基于美国证监会文件设计三类分析任务:单份披露、跨公司对比、长期追踪。
  • 模型在跨时间任务中准确率下降18.60%,暴露幻觉与实体错配问题。
  • 首次量化诊断错误来源,适合金融AI研发和评估人员使用。

随着大语言模型在金融领域的部署增加,对其解析复杂监管披露文件的能力要求越来越高。然而现有基准多聚焦孤立细节,难以反映需整合多份文件、多个报告期和企业实体的专业分析复杂性。此外,这些基准无法区分错误是源于检索失败、生成偏差、领域推理失误,还是对查询或上下文的误解,导致性能瓶颈难以精准定位。为此,我们提出Fin-RATE,一个基于美国证券交易委员会(SEC) filings 的评估基准,模拟金融分析师工作流程,包含三类路径:单份披露内的细节推理、共享主题下的跨企业比较、同一企业跨报告期的纵向追踪。我们对17个主流LLM(涵盖开源、闭源及金融专用模型)在真值上下文与检索增强两种设置下进行评测。结果表明,当任务从单文档推理转向纵向和跨企业分析时,准确率分别下降18.60%和14.35%。这种退化与比较幻觉、时间与实体错配密切相关,并体现在推理质量与事实一致性下降——这些局限性在现有基准中尚未被正式分类或量化。

原文摘要 · Abstract (English)

With the increasing deployment of Large Language Models (LLMs) in the finance domain, LLMs are increasingly expected to parse complex regulatory disclosures. However, existing benchmarks often focus on isolated details, failing to reflect the complexity of professional analysis that requires synthesizing information across multiple documents, reporting periods, and corporate entities. Furthermore, these benchmarks do not disentangle whether errors arise from retrieval failures, generation inaccuracies, domain-specific reasoning mistakes, or misinterpretation of the query or context, making it difficult to precisely diagnose performance bottlenecks. To bridge these gaps, we introduce Fin-RATE, a benchmark built on U.S. Securities and Exchange Commission (SEC) filings and mirroring financial analyst workflows through three pathways: detail-oriented reasoning within individual disclosures, cross-entity comparison under shared topics, and longitudinal tracking of the same firm across reporting periods. We benchmark 17 leading LLMs, spanning open-source, closed-source, and finance-specialized models, under both ground-truth context and retrieval-augmented settings. Results show substantial performance degradation, with accuracy dropping by 18.60% and 14.35% as tasks shift from single-document reasoning to longitudinal and cross-entity analysis. This degradation is associated with increased comparison hallucinations, temporal and entity mismatches, and is further reflected in declines in reasoning quality and factual consistency--limitations that existing benchmarks have yet to formally categorize or quantify.

金融AI评测基准LLM评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。