arXiv:2506.18272cs.CV2025-06中稿 · CODS-COMAD Decembe…被引 1

提出可插拔框架,修复图像解释中漏检与虚构物体问题。

ReFrame: Rectification Framework for Image Explaining Architectures

  • 通过可插入多类模型的可解释框架,修正错误或缺失对象
  • 在图像描述任务中提升完整性81.81%、不一致率降低37.10%
  • 适用于图像描述、视觉问答和基于提示的AI,适合改进现有解释系统

图像解释是深度学习领域的重要研究方向。现有方法常出现幻觉(虚构不存在物体)或遗漏真实物体的问题。本文提出一种可插拔的校正框架 ReFrame,可集成于图像描述、视觉问答(VQA)及基于大语言模型的提示式AI系统,以修正解释中的对象错误。通过基于物体的精确度指标评估,该框架在图像描述任务中使完整性提升81.81%,不一致性降低37.10%;在视觉问答中分别提升9.6%与37.10%;在提示式AI中分别改善0.01%与5.2%。实验表明,其性能显著优于现有最先进方法。

原文摘要 · Abstract (English)

Image explanation has been one of the key research interests in the Deep Learning field. Throughout the years, several approaches have been adopted to explain an input image fed by the user. From detecting an object in a given image to explaining it in human understandable sentence, to having a conversation describing the image, this problem has seen an immense change throughout the years, However, the existing works have been often found to (a) hallucinate objects that do not exist in the image and/or (b) lack identifying the complete set of objects present in the image. In this paper, we propose a novel approach to mitigate these drawbacks of inconsistency and incompleteness of the objects recognized during the image explanation. To enable this, we propose an interpretable framework that can be plugged atop diverse image explaining frameworks including Image Captioning, Visual Question Answering (VQA) and Prompt-based AI using LLMs, thereby enhancing their explanation capabilities by rectifying the incorrect or missing objects. We further measure the efficacy of the rectified explanations generated through our proposed approaches leveraging object based precision metrics, and showcase the improvements in the inconsistency and completeness of image explanations. Quantitatively, the proposed framework is able to improve the explanations over the baseline architectures of Image Captioning (improving the completeness by 81.81% and inconsistency by 37.10%), Visual Question Answering(average of 9.6% and 37.10% in completeness and inconsistency respectively) and Prompt-based AI model (0.01% and 5.2% for completeness and inconsistency respectively) surpassing the current state-of-the-art by a substantial margin.

图像解释可解释性视觉问答大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。