构建首个金融分类体系对齐的多文档审计评测基准,测试大模型在真实财报中的推理能力。
FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMs
- 基于真实XBRL财报构建结构化评测数据集,覆盖三类专业审计任务。
- 13个主流大模型在跨文档一致性推理上表现不佳,平均准确率不足60%。
- 适合研究金融AI、大模型可解释性及合规性验证的学者和工程师。
金融审计超越简单的文本处理,需在大规模披露文件中检测语义、结构和数值不一致。由于财务报告以受会计准则约束的结构化XML格式XBRL提交,审计成为涉及概念对齐、分类体系定义的关系建模以及跨文档一致性的信息抽取与推理问题。尽管大语言模型在孤立的财务任务中表现良好,但其在专业级审计中的能力仍不明确。我们提出FinAuditing,一个基于真实XBRL文件构建的分类体系对齐、结构感知的基准,包含1,102个标注实例,平均每个超过33,000个词元,定义了三项任务:财务语义匹配(FinSM)、财务关系抽取(FinRE)和财务数学推理(FinMR)。对13个先进大模型的评估显示,其在概念检索、分类体系感知的关系建模和一致的跨文档推理方面存在显著差距。研究结果凸显了开发真实、结构感知评测基准的重要性。我们已在GitHub发布评估代码(https://github.com/The-FinAI/FinAuditing),在Hugging Face发布数据集(https://huggingface.co/collections/TheFinAI/finauditing)。该任务目前作为正在进行的公开竞赛SecureFinAI_Contest_2026的官方评测基准。
原文摘要 · Abstract (English)
Going beyond simple text processing, financial auditing requires detecting semantic, structural, and numerical inconsistencies across large-scale disclosures. As financial reports are filed in XBRL, a structured XML format governed by accounting standards, auditing becomes a structured information extraction and reasoning problem involving concept alignment, taxonomy-defined relations, and cross-document consistency. Although large language models (LLMs) show promise on isolated financial tasks, their capability in professional-grade auditing remains unclear. We introduce FinAuditing, a taxonomy-aligned, structure-aware benchmark built from real XBRL filings. It contains 1,102 annotated instances averaging over 33k tokens and defines three tasks: Financial Semantic Matching (FinSM), Financial Relationship Extraction (FinRE), and Financial Mathematical Reasoning (FinMR). Evaluations of 13 state-of-the-art LLMs reveal substantial gaps in concept retrieval, taxonomy-aware relation modeling, and consistent cross-document reasoning. These findings highlight the need for realistic, structure-aware benchmarks. We release the evaluation code at https://github.com/The-FinAI/FinAuditing and the dataset at https://huggingface.co/collections/TheFinAI/finauditing. The task currently serves as the official benchmark of an ongoing public evaluation contest at https://open-finance-lab.github.io/SecureFinAI_Contest_2026/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。