arXiv:2603.19562cs.LGcs.IT2026-03被引 1

揭示视觉对抗脆弱与大模型幻觉的共同几何根源,提出统一理论框架。

Neural Uncertainty Principle: A Unified View of Adversarial Fragility and LLM Hallucination

  • 基于损失诱导态建立神经不确定性原理,揭示输入与梯度的共轭关系。
  • 在近边界状态下,压缩会加剧敏感性分散(对抗脆弱),弱耦合导致生成失控(幻觉)。
  • 设计单向后向探针,可无代价检测幻觉风险,适用于视觉与语言任务。

视觉中的对抗脆弱性与大语言模型中的幻觉传统上被视为独立问题,分别采用特定模态的修复方法。本研究首次揭示二者具有共同的几何起源:输入与其损失梯度是受不可约不确定性约束的共轭可观测量。在损失诱导态下形式化神经不确定性原理(NUP),发现当系统接近该不确定性边界时,进一步压缩必须伴随敏感性分散的增加(即对抗脆弱性);而提示梯度耦合过弱则导致生成过程欠约束(即幻觉)。关键的是,该边界由一个输入-梯度相关通道调节,可通过专门设计的单向后向探针捕捉。在视觉任务中,遮蔽高耦合成分即可提升鲁棒性,无需昂贵的对抗训练;在语言任务中,同一预填充阶段探针可在生成任何答案令牌前检测幻觉风险。因此,NUP将两个看似独立的失效分类转化为共享的不确定性预算视角,并为可靠性分析提供原则性工具。基于此理论,我们提出 ConjMask(遮蔽高贡献输入组件)和 LogitReg(对数概率侧正则化)以在无对抗训练下提升鲁棒性;同时利用探针作为解码无关的风险信号,实现大模型幻觉检测与提示选择。NUP为此类感知与生成任务的边界异常诊断与缓解提供了统一且实用的框架。

原文摘要 · Abstract (English)

Adversarial vulnerability in vision and hallucination in large language models are conventionally viewed as separate problems, each addressed with modality-specific patches. This study first reveals that they share a common geometric origin: the input and its loss gradient are conjugate observables subject to an irreducible uncertainty bound. Formalizing a Neural Uncertainty Principle (NUP) under a loss-induced state, we find that in near-bound regimes, further compression must be accompanied by increased sensitivity dispersion (adversarial fragility), while weak prompt-gradient coupling leaves generation under-constrained (hallucination). Crucially, this bound is modulated by an input-gradient correlation channel, captured by a specifically designed single-backward probe. In vision, masking highly coupled components improves robustness without costly adversarial training; in language, the same prefill-stage probe detects hallucination risk before generating any answer tokens. NUP thus turns two seemingly separate failure taxonomies into a shared uncertainty-budget view and provides a principled lens for reliability analysis. Guided by this NUP theory, we propose ConjMask (masking high-contribution input components) and LogitReg (logit-side regularization) to improve robustness without adversarial training, and use the probe as a decoding-free risk signal for LLMs, enabling hallucination detection and prompt selection. NUP thus provides a unified, practical framework for diagnosing and mitigating boundary anomalies across perception and generation tasks.

神经不确定性对抗脆弱大模型幻觉统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。