arXiv:2607.19259cs.LGcs.AI2026-07中稿 · IJCAI

用大模型融合财报数据与文本,提升财务造假检测泛化能力

Benchmarking Generalization in Financial Statement Fraud Detection: robust evaluation and novel tasks

论文配图:Benchmarking Generalization in Financial Statement Fraud Detection: robust evaluation and novel tasks
图 1 · 摘自论文原文
  • 结合结构化数据与摘要文本,用大模型统一建模
  • 新基准任务下准确率显著优于现有方法
  • 适合金融风控、审计与AI合规研究者参考

财务报表欺诈检测对维护市场诚信至关重要,但面临欺诈手法日益复杂、文本信息利用不足等问题。现有方法多依赖随机数据划分,导致性能评估过于乐观,无法反映真实场景中对新公司或未来时期的有效泛化。为此,我们提出一个基于大语言模型的稳健框架,整合财务报表的结构化数据与管理层讨论与分析(MD&A)的非结构化文本。构建并公开了一个涵盖美国上市公司财务报表、摘要文本及欺诈标签的综合性数据集。设计了更具挑战性的公司隔离式欺诈检测任务(CI-FSFD),实现该任务下的最佳表现,验证了文本信息与鲁棒评估在可靠欺诈检测中的关键价值。

原文摘要 · Abstract (English)

Financial statement fraud detection (FSFD) is crucial for market integrity but faces challenges from increasingly sophisticated schemes and under-utilized textual data in financial reports. Existing methods often rely on random data splits, leading to overoptimistic performance estimates that do not reflect real-world generalization to new companies or future periods. To address this recurring problem with the state of the art, we propose a robust FSFD framework leveraging Large Language Models (LLMs) to integrate both structured financial data and unstructured textual information from financial reports. We provide a more realistic evaluation through a novel and challenging benchmark task called Company-Isolated FSFD (CI-FSFD). We construct and make publicly available a comprehensive U.S. company dataset combining financial statements, summarized MD&A text, and fraud labels. Our approach achieves the best performance on the challenging CI-FSFD task, demonstrating the critical value of textual data and robust evaluation for reliable financial fraud detection.

财务欺诈大模型文本融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。