arXiv:2606.26923cs.CL2026-06被引 1

提出图像文本对错位检测与定位新任务,提升视觉语言模型的可信度。

GAVEL: Grounded Caption Error Verification and Localization

论文配图:GAVEL: Grounded Caption Error Verification and Localization
图 1 · 摘自论文原文
  • 设计新任务GAVEL,同时验证、解释并定位图文不一致问题。
  • 强闭源模型在该任务上表现不佳,表明现有模型仍易幻觉。
  • 首次提供可学习的标注数据,适合研究模型对齐与可解释性的人看。

视觉语言模型常产生幻觉或不一致输出,导致图文内容不匹配。解决此问题不仅需检测错位,还需解释原因并定位视觉证据。本文提出GAVEL(Grounded Caption Error Verification and Localization)任务,联合实现验证、解释与定位。为支持系统评估,构建了对应数据集与基准测试。进一步在人工标注的训练集上训练监督基线模型,检验GAVEL能否提供可学习的监督信号。实验表明,即使强大闭源模型在该任务上也表现欠佳,而监督基线在多个对齐与解释指标上均实现稳定提升。

原文摘要 · Abstract (English)

Vision-language models (VLMs) often produce hallucinated or inconsistent outputs, where text and images are not properly aligned. Addressing this issue requires not only detecting misalignment but also explaining the discrepancy and localizing its visual evidence. We introduce GAVEL (Grounded Caption Error Verification and Localization), a task that jointly addresses verification, explanation, and localization for image-text pairs. To support systematic evaluation, we also present a corresponding dataset and benchmark. We further train a supervised baseline on the human-annotated training split to assess whether GAVEL provides learnable supervision for these abilities. Experiments show that even strong closed-source models struggle on GAVEL, while the supervised baseline yields consistent improvements across grounding and explanation metrics.

视觉语言错误定位模型可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。