arXiv:2606.15949cs.CL2026-06被引 1

构建多文档会计对账基准,评估大模型真实财务整合能力

FinBalance: A Multi-Document Accounting Reconciliation Benchmark

论文配图:FinBalance: A Multi-Document Accounting Reconciliation Benchmark
图 1 · 摘自论文原文
  • 用生成器构建8行业、5难度的源文档包,自动生成凭证与报表
  • 6大模型最终资产负债表准确率最高仅46%,存在26-41个百分点偏差
  • 适合金融合规、财报审计方向研究者,验证模型跨文档推理能力

现有金融NLP基准多评测已生成的文件、表格或提取值。真实会计工作始于更早阶段:需将原始文档对账为引用的会计分录,汇总成资产负债表,并检查矛盾。我们提出FinBalance,一个基于八行业、三种周期类型、五种难度级别的多文档会计对账基准。通过确定性生成器构建人类撰写的业务场景、会计政策、税务/外汇处理、文档结构、干扰项及不一致模板,生成包含分录、资产负债表和23类不一致代码的账本。在710条评估数据上,六款主流LLM最高仅达46%的精确资产负债表准确率。四款模型显示,其报告的资产负债表(BS_exact)与通过账本回放其分录所得结果(BS_recon)间存在26-41个百分点差距。模型常能恢复数值合理的分录,但无法正确关联支持文档,且聚合不一致。引用压力提示几乎未改善文档链接错误,而账本反馈消融实验显著提升报告资产负债表质量,并揭示不一致检测的权衡。专家财务评审确认了基准设计与标签的有效性。

原文摘要 · Abstract (English)

Existing financial-NLP benchmarks mostly evaluate prepared artifacts such as filings, tables, or extracted values. Real accounting begins earlier: source documents must be reconciled into cited journal entries, aggregated into a balance sheet, and checked for contradictions. We introduce FinBalance, a multi-document accounting reconciliation benchmark built from source-document bundles across eight industries, three period types, and five difficulty levels. Human-authored business scenarios, accounting policies, tax/FX treatments, document schemas, distractors, and inconsistency templates are composed by a deterministic generator whose ledger produces journal entries,balance sheets, and 23 inconsistency-code labels. On a 710-record evaluation split, six contemporary LLMs reach at most 46% exact final-balance-sheet accuracy. Four models show a 26-41 pp gap between BS_exact, the model's reported balance sheet, and BS_recon, the balance sheet obtained by replaying its entries through our ledger. Models often recover numerically plausible entries but fail to bind them to supporting documents and aggregate them consistently. Citation-pressure prompting barely changes document-linking errors, while ledger-feedback ablations substantially improve reported balance sheets and expose inconsistency-detection trade-offs. Expert finance reviewers validate the benchmark design and labels.

会计对账多文档金融NLP大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。