提出视觉可靠性度量,可检测手语翻译中模型是否凭空编造。
Grounding or Guessing? Visual Signals for Detecting Hallucinations in Sign Language Translation
- 通过视频遮蔽与反事实对比,量化解码器对视觉信息的依赖程度。
- 可靠性分数能准确预测幻觉率,且在不同数据集和模型间通用。
- 适用于无词素标注的模型,帮助识别哪些词是猜的、哪些是基于视觉的。
幻觉问题——模型生成与视觉证据不符的流畅文本——是视觉语言模型的主要缺陷,尤其在手语翻译(SLT)中尤为关键。由于手语意义高度依赖视频内容,无词素(gloss-free)模型更易出错,因其直接将连续手势映射为自然语言,缺乏中间词素对齐的监督。本文认为幻觉源于模型依赖语言先验而非视觉输入。为此,提出一种基于标记级别的可靠性度量,综合特征敏感性(视频遮蔽时内部变化)与反事实信号(正常与扰动视频下的概率差异),聚合为句子级可靠性分数,可简洁解释视觉对齐程度。在两个主流SLT基准(PHOENIX-2014T 和 CSL-Daily)上评估,结果表明该度量能有效预测幻觉率,跨数据集与架构泛化良好,并在视觉退化下下降。定性分析显示,该度量可区分有依据与凭空猜测的词,无需参考文本即可估计风险;结合文本置信度、困惑度或熵值后,进一步提升幻觉检测能力。研究证明可靠性度量是诊断手语翻译中幻觉的实用工具,为多模态生成中的鲁棒检测奠定基础。
原文摘要 · Abstract (English)
Hallucination, where models generate fluent text unsupported by visual evidence, remains a major flaw in vision-language models and is particularly critical in sign language translation (SLT). In SLT, meaning depends on precise grounding in video, and gloss-free models are especially vulnerable because they map continuous signer movements directly into natural language without intermediate gloss supervision that serves as alignment. We argue that hallucinations arise when models rely on language priors rather than visual input. To capture this, we propose a token-level reliability measure that quantifies how much the decoder uses visual information. Our method combines feature-based sensitivity, which measures internal changes when video is masked, with counterfactual signals, which capture probability differences between clean and altered video inputs. These signals are aggregated into a sentence-level reliability score, providing a compact and interpretable measure of visual grounding. We evaluate the proposed measure on two SLT benchmarks (PHOENIX-2014T and CSL-Daily) with both gloss-based and gloss-free models. Our results show that reliability predicts hallucination rates, generalizes across datasets and architectures, and decreases under visual degradations. Beyond these quantitative trends, we also find that reliability distinguishes grounded tokens from guessed ones, allowing risk estimation without references; when combined with text-based signals (confidence, perplexity, or entropy), it further improves hallucination risk estimation. Qualitative analysis highlights why gloss-free models are more susceptible to hallucinations. Taken together, our findings establish reliability as a practical and reusable tool for diagnosing hallucinations in SLT, and lay the groundwork for more robust hallucination detection in multimodal generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。