arXiv:2512.03107cs.LGcs.CL2025-12被引 2

用信息论方法检测金融领域大模型幻觉,准确率提升92%

Detecting AI Hallucinations in Finance: An Information-Theoretic Method Cuts Hallucination Rate by 92%

  • 通过语义熵与证据容量的不匹配度衡量幻觉
  • 在200个样本上实现0.89的AUC和0.90的精确率
  • 适合关注大模型可信性的金融与AI安全研究者

大语言模型常生成流畅但无依据的内容(幻觉),限制其在高风险领域的应用。本文提出ECLIPSE框架,将幻觉视为模型语义熵与可用证据容量之间的不匹配。通过多样本聚类估计熵,并引入新的困惑度分解来衡量模型对检索证据的使用情况。理论上证明,在温和条件下,该熵-容量目标函数是严格凸的,具有唯一稳定最优解。在包含200个带合成幻觉样本的金融问答数据集上,使用GPT-3.5-turbo评估,ECLIPSE达到0.89的ROC AUC和0.90的平均精度,显著优于仅依赖语义熵的基线(AUC 0.50)。对Claude-3-Haiku的消融实验显示,当缺乏逐标记概率时,AUC降至0.59,系数幅度下降95%,表明ECLIPSE依赖校准的词元级不确定性。困惑度分解特征的系数最大,证实证据利用是幻觉检测的核心。本工作为受控机制研究,跨领域及真实幻觉的广泛验证仍需后续工作。

原文摘要 · Abstract (English)

Large language models (LLMs) produce fluent but unsupported answers - hallucinations - limiting safe deployment in high-stakes domains. We propose ECLIPSE, a framework that treats hallucination as a mismatch between a model's semantic entropy and the capacity of available evidence. We combine entropy estimation via multi-sample clustering with a novel perplexity decomposition that measures how models use retrieved evidence. We prove that under mild conditions, the resulting entropy-capacity objective is strictly convex with a unique stable optimum. We evaluate on a controlled financial question answering dataset with GPT-3.5-turbo (n=200 balanced samples with synthetic hallucinations), where ECLIPSE achieves ROC AUC of 0.89 and average precision of 0.90, substantially outperforming a semantic entropy-only baseline (AUC 0.50). A controlled ablation with Claude-3-Haiku, which lacks token-level log probabilities, shows AUC dropping to 0.59 with coefficient magnitudes decreasing by 95% - demonstrating that ECLIPSE is a logprob-native mechanism whose effectiveness depends on calibrated token-level uncertainties. The perplexity decomposition features exhibit the largest learned coefficients, confirming that evidence utilization is central to hallucination detection. We position this work as a controlled mechanism study; broader validation across domains and naturally occurring hallucinations remains future work.

大模型幻觉信息论金融AI可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。