arXiv:2511.21756cs.CLcs.CE2025-11

发现金融大模型算术幻觉根源,定位并抑制关键错误电路。

Dissecting the Ledger: Locating and Suppressing "Liar Circuits" in Financial Large Language Models

  • 通过因果追踪定位算术推理中的中间计算与最终决策电路。
  • 抑制第46层使幻觉输出置信度下降81.8%,验证机制有效性。
  • 该层特征可泛化至未见金融话题,准确率达98%。

大型语言模型在金融领域应用日益广泛,但在执行算术运算时存在特定且可复现的幻觉问题。现有缓解策略多将模型视为黑箱。本文提出一种机制性检测方法,通过对GPT-2 XL在ConvFinQA基准上的因果追踪,揭示算术推理的双阶段机制:中层(L12-L30)为分布式计算暂存区,晚期层(特别是第46层)为决定性聚合电路。通过消融实验验证,抑制第46层可使模型对幻觉输出的置信度降低81.8%。此外,基于该层训练的线性探测器在未见金融主题上达到98%准确率,表明算术欺骗具有普遍几何结构。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly deployed in high-stakes financial domains, yet they suffer from specific, reproducible hallucinations when performing arithmetic operations. Current mitigation strategies often treat the model as a black box. In this work, we propose a mechanistic approach to intrinsic hallucination detection. By applying Causal Tracing to the GPT-2 XL architecture on the ConvFinQA benchmark, we identify a dual-stage mechanism for arithmetic reasoning: a distributed computational scratchpad in middle layers (L12-L30) and a decisive aggregation circuit in late layers (specifically Layer 46). We verify this mechanism via an ablation study, demonstrating that suppressing Layer 46 reduces the model's confidence in hallucinatory outputs by 81.8%. Furthermore, we demonstrate that a linear probe trained on this layer generalizes to unseen financial topics with 98% accuracy, suggesting a universal geometry of arithmetic deception.

大模型幻觉金融AI因果追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。