arXiv:2412.16232cs.CVcs.AI2024-12AAAI被引 14

提出可修正的视觉蕴含任务,让模型动态调整图像与文本关系判断。

Defeasible Visual Entailment: Benchmark, Evaluator, and Reward-Driven Optimization

  • 引入可修正视觉蕴含(DVE)新任务,支持更新修正初始推理
  • 设计新型评估器,通过对比学习捕捉蕴含强度变化,提升评估精度
  • 提出奖励驱动优化方法,改善多模态模型生成更新的质量

我们提出一种新任务——可修正视觉蕴含(Defeasible Visual Entailment, DVE),旨在根据附加更新信息,动态调整图像前提与文本假设之间的蕴含关系。尽管该概念在自然语言推理中已成熟,但在视觉蕴含领域仍属空白。DVE使模型能修正初始判断,从而提升误导信息检测、视觉问答及自动驾驶决策等应用的准确性和可靠性。现有评估指标无法有效捕捉更新带来的蕴含关系变化。为此,我们设计了一种新型推理感知评估器,结合成对对比学习与类别信息学习,精准捕捉更新引发的蕴含强度变化。此外,提出一种奖励驱动的更新优化方法,进一步提升多模态模型生成更新的质量。实验验证了评估器与优化方法的有效性。

原文摘要 · Abstract (English)

We introduce a new task called Defeasible Visual Entailment (DVE), where the goal is to allow the modification of the entailment relationship between an image premise and a text hypothesis based on an additional update. While this concept is well-established in Natural Language Inference, it remains unexplored in visual entailment. At a high level, DVE enables models to refine their initial interpretations, leading to improved accuracy and reliability in various applications such as detecting misleading information in images, enhancing visual question answering, and refining decision-making processes in autonomous systems. Existing metrics do not adequately capture the change in the entailment relationship brought by updates. To address this, we propose a novel inference-aware evaluator designed to capture changes in entailment strength induced by updates, using pairwise contrastive learning and categorical information learning. Additionally, we introduce a reward-driven update optimization method to further enhance the quality of updates generated by multimodal models. Experimental results demonstrate the effectiveness of our proposed evaluator and optimization method.

视觉蕴含多模态推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。