arXiv:2503.04457cs.CVcs.AI2025-03被引 6

通过时间连接提升视觉语言模型逻辑一致性,有效减少幻觉。

TPC: Cross-Temporal Prediction Connection for Vision-Language Model Hallucination Reduction

  • 跨时间步连接预测结果,增强逻辑连续性
  • 在多个基准上显著降低幻觉率,提升生成准确性
  • 适合需要高可靠性的视觉描述与开放生成场景

视觉语言模型(VLMs)在多项任务中取得显著进展,得益于大语言模型(LLMs)的强大能力。然而,模型常因过度依赖语言先验而产生幻觉,即错误地描述图像中不存在的物体或属性,严重影响高风险应用中的可靠性。本文观察到对数输出的连续性一致性特征,提出一种简单高效的跨时间预测连接(TPC)方法,通过在不同时间步间连接对数输出,增强语义一致性,提升信息流动与生成连贯性,有效缓解幻觉问题。大量实验表明,TPC在准确性和效率上均优于现有主流方法,在开放式文本生成任务中保持鲁棒性。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have achieved remarkable advancements, capitalizing on the impressive capabilities of large language models (LLMs) across diverse tasks. Despite this, a critical challenge known as hallucination occurs when models overconfidently describe objects or attributes absent from the image, a problem exacerbated by the tendency of VLMs to rely on linguistic priors. This limitation reduces model reliability in high-stakes applications. In this work, we have observed the characteristic of logits' continuity consistency enhancement and introduced a straightforward and efficient method, Cross-Temporal Prediction Connection (TPC), designed to enhance the semantic consistency of logits by connecting them temporally across timesteps. TPC amplifies information flow and improves coherence, effectively reducing hallucination. Extensive experiments show that TPC surpasses existing representatives, delivering superior performance in both accuracy and efficiency while maintaining robustness in open-ended text generation tasks.

视觉语言模型幻觉抑制生成一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。