构建金融表格幻觉评估框架,精准检测模型数值错误
FAITH: A Framework for Assessing Intrinsic Tabular Hallucinations in Finance
- 设计上下文感知的掩码预测任务,自动构建金融数据评估集
- 基于标普500年报生成新数据集,覆盖真实金融场景中的数值错误
- 揭示主流大模型在财务表格中的内在幻觉模式,适合金融AI研发者使用
幻觉仍是将大语言模型(LLMs)应用于金融领域的关键挑战。从表格数据中准确提取并精确计算对可靠金融分析至关重要,因为即使微小的数值错误也可能破坏决策并导致合规风险。金融应用具有独特需求,常依赖上下文相关的、数值型且专有的表格数据,而现有幻觉评估基准很少涵盖此类场景。本文提出一个严谨且可扩展的框架,用于评估金融领域大模型的内在幻觉,将其建模为在真实金融文档上的上下文感知掩码跨度预测任务。主要贡献包括:(1)一种新颖的自动化数据集构建范式,采用掩码策略;(2)源自标普500年度报告的新幻觉评估数据集;(3)对先进大模型在金融表格数据上内在幻觉模式的全面评估。本工作为内部大模型评估提供了可靠方法,是构建更可信、更可靠的金融生成式AI系统的关键一步。
原文摘要 · Abstract (English)
Hallucination remains a critical challenge for deploying Large Language Models (LLMs) in finance. Accurate extraction and precise calculation from tabular data are essential for reliable financial analysis, since even minor numerical errors can undermine decision-making and regulatory compliance. Financial applications have unique requirements, often relying on context-dependent, numerical, and proprietary tabular data that existing hallucination benchmarks rarely capture. In this study, we develop a rigorous and scalable framework for evaluating intrinsic hallucinations in financial LLMs, conceptualized as a context-aware masked span prediction task over real-world financial documents. Our main contributions are: (1) a novel, automated dataset creation paradigm using a masking strategy; (2) a new hallucination evaluation dataset derived from S&P 500 annual reports; and (3) a comprehensive evaluation of intrinsic hallucination patterns in state-of-the-art LLMs on financial tabular data. Our work provides a robust methodology for in-house LLM evaluation and serves as a critical step toward building more trustworthy and reliable financial Generative AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。