arXiv:2502.14908cs.CVcs.AI2025-02被引 1

提出新方法检测视觉语言模型在知识冲突下的幻觉问题。

SegSub: Evaluating Robustness to Knowledge Conflicts and Hallucinations in Vision-Language Models

  • 设计图像扰动框架,精准测试模型对跨模态冲突的抗性。
  • 发现模型识别反事实条件准确率低于30%,解决来源冲突不足1%。
  • 适合关注多模态系统可靠性与安全性的研究者参考。

视觉语言模型虽具备复杂多模态推理能力,但在面对知识冲突时易产生幻觉,限制其在信息敏感场景的应用。现有研究多聚焦单模态模型的鲁棒性,而多模态领域缺乏对跨模态知识冲突的系统性分析。本文提出 exttt{SegSub}框架,通过针对性图像扰动评估视觉语言模型对知识冲突的韧性。分析显示:模型对参数冲突具有较强鲁棒性(20%遵循率),但对反事实条件识别准确率低于30%,来源冲突解决准确率不足1%。上下文丰富度与幻觉率呈负相关(r = -0.368, p = 0.003),揭示高风险图像特征。基于基准数据集的微调实验表明,模型在知识冲突检测上获得提升,为构建信息敏感环境下的抗幻觉多模态系统奠定基础。

原文摘要 · Abstract (English)

Vision language models (VLM) demonstrate sophisticated multimodal reasoning yet are prone to hallucination when confronted with knowledge conflicts, impeding their deployment in information-sensitive contexts. While existing research addresses robustness in unimodal models, the multimodal domain lacks systematic investigation of cross-modal knowledge conflicts. This research introduces \segsub, a framework for applying targeted image perturbations to investigate VLM resilience against knowledge conflicts. Our analysis reveals distinct vulnerability patterns: while VLMs are robust to parametric conflicts (20% adherence rates), they exhibit significant weaknesses in identifying counterfactual conditions (<30% accuracy) and resolving source conflicts (<1% accuracy). Correlations between contextual richness and hallucination rate (r = -0.368, p = 0.003) reveal the kinds of images that are likely to cause hallucinations. Through targeted fine-tuning on our benchmark dataset, we demonstrate improvements in VLM knowledge conflict detection, establishing a foundation for developing hallucination-resilient multimodal systems in information-sensitive environments.

视觉语言模型幻觉检测多模态鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。