arXiv:2608.07400cs.AIcs.DB2026-08

构建金融问答新基准,确保答案与原始披露证据精准匹配。

FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings

  • 以真实财报为依据,要求系统定位正确的企业、时间与披露上下文。
  • 1185个问题涵盖22家公司财报,含精心设计的干扰项,挑战模型识别力。
  • 适合研究金融AI、可解释性或需要高可信度问答系统的团队使用。

金融问答通常仅评估答案正确性,但财报中看似合理甚至数值正确的回答可能基于错误证据。同一公司不同期间或相似公司间存在大量相似事实与披露内容。FinRank针对这一依赖证据来源的检索难题,要求系统准确识别目标实体、报告期及披露语境。该基准包含22家公司的10-K与10-Q文件中人工撰写的1185条问答记录,每条均配有参考答案、黄金支持段落,以及从同一文件、不同报告期及可比公司中提取的精心设计的难负样本。该基准独立评估段落检索、重排序及难负样本区分能力。基线结果显示任务难度极高:70亿参数指令微调嵌入模型在合并证据库上仅达44.8% Recall@10;小于十亿参数编码器相较BM25提升不超过3.5点;金融适配嵌入模型反而落后BM25 9.7点;当随机负样本替换为定制难负样本时,成对准确率下降13.0至20.5个百分点。FinRank为开发不仅准确且证据可靠的金融问答系统提供了一个以证据为核心的基准。

原文摘要 · Abstract (English)

Financial question answering is typically evaluated by answer correctness, yet in SEC filings a plausible and even numerically correct answer can be grounded in the wrong evidence. Similar facts and disclosures recur across sections of a filing, across reporting periods of the same firm, and across comparable firms. FinRank targets this provenance-sensitive retrieval problem by requiring systems to identify evidence for the intended entity, reporting period, and disclosure context. The benchmark contains 1185 manually authored question-answer records over the 10-K and 10-Q filings of 22 companies. Each record includes a reference answer, gold supporting passages, and hand-curated hard negatives drawn from confusable passages within filings, across reporting periods, and across comparable firms. FinRank evaluates passage retrieval, reranking, and hard-negative discrimination as separately measured tasks. Baseline results demonstrate the difficulty of this setting: among the evaluated systems, even a 7B instruction-tuned embedder reaches only 44.8% Recall@10 on the pooled evidence corpus; sub-billion-parameter encoders gain at most 3.5 points over BM25, a finance-adapted embedder trails BM25 by 9.7 points, and pairwise accuracy falls by 13.0-20.5 percentage points when random negatives are replaced with the curated hard negatives. FinRank provides an evidence-first benchmark for developing financial question answering systems that are not only accurate but also grounded in the correct disclosure.

金融AI证据溯源问答系统财报分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。