arXiv:2511.08596cs.CLcs.AI2025-11被引 1

用自生成谎言测试大模型抗干扰能力,发现不同模型可靠性差异巨大

What About the Scene with the Hitler Reference? HAUNT: A Framework to Probe LLMs' Self-consistency Via Adversarial Nudge

  • 让模型先造真话假话,再自己验证,最后用假话骗它
  • GPT/Grok中等抗骗,Gemini/DeepSeek易被误导,Claude最稳
  • 适合关注大模型事实性的人看,尤其信息检索场景

幻觉是大语言模型在高风险领域实际部署中的关键挑战。本文提出一种框架,用于在对抗性诱导下对大模型的事实一致性进行压力测试。该框架包含三个步骤:首先指令模型生成与特定封闭领域一致的真话和假话;其次指令模型对同一组陈述进行真假判断;最后测试模型对自身生成并验证过的假话的鲁棒性。我们在两个流行封闭领域(电影、小说)上对五种知名专有大模型进行了广泛评估,结果显示模型对对抗性诱导的敏感度差异显著: exttt{Claude} 展现出强韧性, exttt{GPT} 和 exttt{Grok} 表现中等,而 exttt{Gemini} 与 exttt{DeepSeek} 则表现出弱韧性。鉴于越来越多用户依赖大模型获取信息,这些发现令人警觉。

原文摘要 · Abstract (English)

Hallucinations pose a critical challenge to the real-world deployment of large language models (LLMs) in high-stakes domains. In this paper, we present a framework for stress testing factual fidelity in LLMs in the presence of adversarial nudge. Our framework consists of three steps. In the first step, we instruct the LLM to produce sets of truths and lies consistent with the closed domain in question. In the next step, we instruct the LLM to verify the same set of assertions as truths and lies consistent with the same closed domain. In the final step, we test the robustness of the LLM against the lies generated (and verified) by itself. Our extensive evaluation, conducted using five widely known proprietary LLMs across two closed domains of popular movies and novels, reveals a wide range of susceptibility to adversarial nudges: \texttt{Claude} exhibits strong resilience, \texttt{GPT} and \texttt{Grok} demonstrate moderate resilience, while \texttt{Gemini} and \texttt{DeepSeek} show weak resilience. Considering that a large population is increasingly using LLMs for information seeking, our findings raise alarm.

大模型幻觉对抗测试事实一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。