arXiv:2606.09859cs.LGcs.AI2026-06

提出几何感知的解码方法,抑制幻觉同时保持语义结构稳定

Mitigating Manifold Departure: Uncertainty-Aware Subspace Rectification for Trustworthy MLLM Decoding

论文配图:Mitigating Manifold Departure: Uncertainty-Aware Subspace Rectification for Trustworthy MLLM Decoding
图 1 · 摘自论文原文
  • 基于SVD构建语言先验子空间,动态投影并选择性抑制噪声成分
  • 在POPE和CHAIR上显著降低幻觉率,且不损失文本连贯性
  • 适合追求高可信度多模态生成的应用场景

MLLMs常因过度依赖语言先验而生成与视觉输入不符的物体。现有无训练解码策略通过惩罚语言先验来缓解此问题,但忽略了语言先验的双重性:其在与视觉证据对齐时有益,偏离时有害。盲目压制先验会破坏模型的语义流形,导致性能下降,我们称此为「流形偏移」。为此,本文提出流形引导自适应投影(MGAP),一种几何感知的无训练解码方法,可在抑制幻觉的同时保持表征结构。MGAP首先通过SVD从盲隐藏状态构建语言先验子空间;解码过程中,将每个多模态隐藏状态投影至该子空间,并应用一致性门控机制,仅自适应地衰减投影后的先验分量,实现子空间选择性更新,从而保留正交的语义成分。在POPE和CHAIR上的大量实验表明,MGAP优于现有解码基线,在不牺牲连贯性的前提下实现更强的幻觉抑制。

原文摘要 · Abstract (English)

MLLMs frequently hallucinate objects inconsistent with visual inputs. This issue is typically attributed to the over-reliance on language priors, which can override the visual context. Recent training-free decoding strategies address this by penalizing language priors. However, these methods overlook the dual nature of language priors, where they can be both helpful and harmful depending on the alignment with visual evidence. In particular, blindly suppressing language priors often disrupts the model's semantic manifold, leading to performance degradation, a phenomenon we term Manifold Departure. To address this, we propose Manifold-Guided Adaptive Projection (MGAP), a geometry-aware, training-free decoding method that mitigates hallucinations while preserving representation structure. MGAP first constructs a language-prior subspace from blind hidden states via SVD. During decoding, MGAP projects each multimodal hidden state onto this subspace and applies a consistency-aware gate to adaptively attenuate only the projected prior component, yielding a subspace-selective update that largely preserves the orthogonal semantic components. Extensive experiments on POPE and CHAIR show that MGAP outperforms prior decoding baselines, achieving stronger hallucination suppression without sacrificing coherence.

多模态生成幻觉抑制流形学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。