arXiv:2603.11339cs.AIcs.CE2026-03被引 6

评测大模型在真实财报中联合推理会计规则的能力。

FinRule-Bench: A Benchmark for Joint Reasoning over Financial Tables and Principles

  • 构建真实财报与人工标注准则的配对数据集,支持多任务评估。
  • 三类任务逐步提升难度:验证单条规则、识别违规规则、定位多处错误。
  • 发现模型在复杂诊断任务中表现显著下降,暴露推理局限性。

大型语言模型(LLMs)在金融分析中应用日益广泛,但其在明确会计准则下审计结构化财务报表的能力尚未充分探索。现有基准主要评估问答、数值推理或合成数据中的异常检测,难以判断模型能否在正确财务报表上可靠验证或定位规则合规性。我们提出FinRule-Bench,一个面向真实世界财务表格的规则型金融推理诊断完整性评估基准。该基准将真实财务报表与人工标注的会计准则配对,覆盖四大典型报表类型:资产负债表、现金流量表、利润表和权益变动表。定义三项审计任务,要求逐步增强推理能力:(i) 规则验证,测试单条原则的合规性;(ii) 规则识别,从给定规则集中选出违反的原则;(iii) 联合规则诊断,需在记录层面检测并定位多个同时存在的违规。我们在零样本和少样本提示下评估LLMs,并引入因果-反事实推理协议,强制决策、解释与反事实判断之间的一致性。在各类任务和报表类型中,模型在孤立规则验证上表现良好,但在规则区分和多违规诊断任务中性能急剧下降。FinRule-Bench为研究高风险金融分析中基于规则的推理、诊断覆盖率及模型失效模式提供了可复现的测试平台。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly applied to financial analysis, yet their ability to audit structured financial statements under explicit accounting principles remains poorly explored. Existing benchmarks primarily evaluate question answering, numerical reasoning, or anomaly detection on synthetically corrupted data, making it unclear whether models can reliably verify or localize rule compliance on correct financial statements. We introduce FinRule-Bench, a benchmark for evaluating diagnostic completeness in rule-based financial reasoning over real-world financial tables. FinRule-Bench pairs ground-truth financial statements with explicit, human-curated accounting principles and spans four canonical statement types: Balance Sheets, Cash Flow Statements, Income Statements, and Statements of Equity. The benchmark defines three auditing tasks that require progressively stronger reasoning capabilities: (i) rule verification, which tests compliance with a single principle; (ii) rule identification, which requires selecting the violated principle from a provided rule set; and (iii) joint rule diagnosis, which requires detecting and localizing multiple simultaneous violations at the record level. We evaluate LLMs under zero-shot and few-shot prompting, and introduce a causal-counterfactual reasoning protocol that enforces consistency between decisions, explanations, and counterfactual judgments. Across tasks and statement types, we find that while models perform well on isolated rule verification, performance degrades sharply for rule discrimination and multi-violation diagnosis. FinRule-Bench provides a principled and reproducible testbed for studying rule-governed reasoning, diagnostic coverage, and failure modes of LLMs in high-stakes financial analysis.

金融分析规则推理大模型评测财务审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。