金融大模型在领域微调后数值幻觉激增,根源是数值克制力下降而非推理不足。
When Financial Fine-tuning Fails: A Three-Level Detectability Analysis of Numerical Hallucination in Domain-Adapted Language Models

- 构建三层次幻觉检测分类,区分显性、隐性显式和隐性隐式幻觉。
- 领域微调使显性幻觉率从5.4%飙升至82.5%,增强数值训练后达98%。
- 发现模板注入是核心幻觉机制,建议部署时加入事实核查或拒答机制。
金融大语言模型在报告摘要任务中日益广泛应用,但数值幻觉带来重大实践风险。现有研究常归因于数值推理不足,但未在受控微调条件下系统验证。本文在三种模型变体上开展成本可控的实验:基础指令微调模型、领域语言适配模型(FT-A)和数值增强领域模型(FT-A+B+C)。提出三层次可检测性分类:显性幻觉(货币单位虚构)、隐性显式幻觉(专业惯例数值)和隐性隐式幻觉(无依据定量陈述)。结果表明,领域微调显著降低所有层级的数值克制力:基础模型幻觉率仅5.4%,而FT-A达82.5%,FT-A+B+C高达98%。出乎意料的是,数值监督反而加剧幻觉。识别出模板注入——即无视输入内容插入记忆中的标准值——是微调模型的主要幻觉机制。研究证明,金融摘要中的数值幻觉源于领域适配导致的数值克制力退化,而非推理能力不足。建议评估应覆盖所有可检测层级,部署应包含基于事实生成或拒绝生成的机制。
原文摘要 · Abstract (English)
Financial large language models are increasingly deployed for summarization of reports and disclosures, where numerical hallucination poses significant practical risks. While prior work often attributes such hallucination to insufficient numerical reasoning, this assumption has not been systematically tested under controlled fine-tuning settings. In this paper, we conduct a cost-effective, controlled study of numerical hallucination in financial summarization across three model variants: a base instruction-tuned model, a domain language-adapted model (FT-A), and a numeracy-enhanced domain model (FT-A+B+C). We introduce a three-level detectability taxonomy distinguishing between overt hallucination (currency-denominated fabrication), covert-explicit hallucination (professional-convention numbers), and covert-implicit hallucination (ungrounded quantitative claims). Our results reveal that domain fine-tuning substantially degrades numerical restraint at all detectability levels. While the Base model maintains near-zero hallucination rates (5.4\%), FT-A exhibits 82.5\% overt hallucination and FT-A+B+C reaches 98\%. Contrary to intuition, numeracy supervision amplifies rather than mitigates hallucination across all levels. We identify template injection---the insertion of memorized canonical values regardless of input content---as a primary hallucination mechanism in fine-tuned models. These findings demonstrate that numerical hallucination in financial summarization is driven by the degradation of numerical restraint through domain adaptation, not by insufficient numerical reasoning. We recommend that evaluation protocols assess hallucination across all detectability levels and that deployment practices include explicit mechanisms for grounding-aware generation or abstention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。