arXiv:2608.18534cs.AIcs.IR2026-08

金融AI系统诊断能力被证据获取拖累,新基准揭示检索才是关键瓶颈。

FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems

  • 构建2250个合成对账案例,分离证据检索与推理评估。
  • 检索准确率提升使整体正确率从2.05%升至72.44%。
  • 模型正确判断常依赖不完整证据,需关注可审计的证据链。

大语言模型在金融操作中日益普及,但其推理表现可能受是否获得正确证据影响。在应付账款对账中,诊断所需证据分散于发票、采购单、审批、分配、付款、账簿记录和银行流水之间,通过交易关系连接而非文本相似性。因此端到端准确率易混淆证据获取与推理质量。我们提出FinRCA-Bench,一个包含2,250个应付账款对账案例的确定性合成基准,覆盖14张运营表,含1,500个注入故障(15类因果)和750个合法或难负样本。根因标签与记录级证据合约对模型隐藏,实现检索独立评估。对比规则/SQL、经典机器学习、密集语义检索、确定性关系扩展及类型化溯源图检索(TPGR),规则/SQL达到84.97%的保留精确度,经典机器学习达95.44%。固定推理模型与提示设置,仅更换检索方式,宏观所需记录召回率从0.83%升至77.70%,16类精确度从2.05%升至72.44%。结构化检索失败是推理失败的6倍以上(95比15);254次正确预测发生在证据不全情况下,严格返回证据合约准确率仅5.72%。结果表明,检索架构显著影响系统性能,而正确根因标签难以代表可审计诊断。

原文摘要 · Abstract (English)

Large language models are increasingly used to support financial operations, but their apparent reasoning performance can depend on whether they receive the right evidence. In financial reconciliation, the evidence needed for diagnosis is distributed across invoices, purchase orders, approvals, allocations, payments, ledger entries, and bank activity, linked by transactional relationships rather than textual similarity. End-to-end accuracy can therefore conflate evidence access with reasoning quality. We introduce FinRCA-Bench, a deterministic synthetic benchmark of 2,250 accounts-payable-to-bank reconciliation cases spanning 14 operational tables, including 1,500 injected failures across 15 causal categories and 750 legitimate or hard-negative cases. Root-cause labels and record-level evidence contracts are hidden from the model, allowing retrieval to be evaluated independently of answer correctness. We compare Rules/SQL, classical machine learning, dense semantic retrieval, deterministic relational expansion, and Typed Provenance Graph Retrieval (TPGR), a typed traversal restricted to persisted transaction relationships. Rules/SQL reaches 84.97% held-out exact accuracy and classical ML reaches 95.44%. Holding the reasoning model, prompt, and generation settings fixed while changing only retrieval increases macro required-record recall from 0.83% to 77.70% and exact 16-class accuracy from 2.05% to 72.44%. Structural retrieval failures outnumber reasoning failures with sufficient retrieval by 95 to 15; 254 correct predictions occur despite incomplete retrieval, and strict returned-evidence contract accuracy is only 5.72%. On FinRCA-Bench, retrieval architecture strongly shapes observed AI-system performance, and a correct root-cause label is a weak proxy for an auditable diagnosis.

金融AI证据检索对账系统可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。