arXiv:2412.07801cs.CVcs.AI2024-12中稿 · ACM MM 2024被引 12

让AI学会像老师一样解释错误,提升视觉推理能力

Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor

  • 用GPT-4生成可解释的纠错反馈数据集
  • 新模型在错误识别和原因说明上显著优于现有模型
  • 适合关注可解释性与教学型AI的研究者

大型多模态模型(LMMs)在视觉常识推理(VCR)任务中表现优异,该任务要求基于图像中的视觉常识回答多选题。然而,当前LMMs在发现并纠正选项中潜在视觉常识错误方面仍研究不足。受人类教师通过设计挑战性干扰项来检测学生理解程度并引导其纠错行为的启发,我们首次探索了让LMM模拟这一纠错过程。为此,我们采用GPT-4作为“教师”,构建了可解释反馈数据集VCR-DF,用于评估LMM识别误解并阐明错误原因的能力。此外,我们提出一种基于LMM的教学专家指导反馈生成模型(PEIFG),通过可学习的专家提示和多模态指令引导反馈生成。实验表明,PEIFG显著优于现有LMMs。我们认为该基准为评估LMM能力提供了新方向。

原文摘要 · Abstract (English)

Large multimodal models (LMMs) have shown remarkable performance in the visual commonsense reasoning (VCR) task, which aims to answer a multiple-choice question based on visual commonsense within an image. However, the ability of LMMs to correct potential visual commonsense errors in the distractor upon their occurrence is yet under-explored. Drawing inspiration from how a human teacher crafts challenging distractors to test students' comprehension of the concepts or skills and assists them in identifying and correcting errors toward the answer, we are the pioneering research for LMMs to simulate this error correction process. To this end, we employ GPT-4 as a ``teacher'' to collect the explainable feedback dataset VCR-DF for error correction, which serves as a benchmark to evaluate the ability of LMMs to identify misconceptions and clarify reasons behind the error in VCR distractors toward final answers. In addition, we propose an LMM-based Pedagogical Expert Instructed Feedback Generation (PEIFG) model to incorporate the learnable expert prompts and multimodal instruction as guidance for feedback generation. Experimental results show that our PEIFG significantly outperforms existing LMMs. We believe that our benchmark provides a new direction for evaluating the capabilities of LMMs.

视觉推理可解释性教学型AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。