让机器检测操作错误时自解释,提升决策透明度与可信度。
Transparent and Coherent Procedural Mistake Detection
- 用视觉语言模型生成带理由的判断过程,实现可解释的错误检测。
- 引入自然语言推理模型评估理由连贯性,提升判断质量。
- 适合关注AI可解释性与人机协作的研究者和开发者。
程序错误检测(PMD)是通过第一人称视频判断用户是否成功执行了由步骤文本指定的任务,是一项极具挑战性的任务。尽管近期有诸多努力,但现有模型在真实场景下的表现仍不可靠,且其决策过程缺乏透明性。为此,我们重新定义PMD任务,要求模型生成视觉自对话式推理过程以支持判断。基于单帧图像构建合适的基准数据集,并利用近期视觉-语言模型(VLMs)强大的图像理解能力,我们提出一种可解释的新范式。为评估生成理由的连贯性,我们设计两个基于自然语言推理(NLI)的自动化指标。实验表明,虽直接使用VLM性能有限,但结合这些指标改进推理与微调方法后,模型的准确性、理由连贯性与效率均得到提升。多维度评估指标可视化常见问题,指明未来优化方向。
原文摘要 · Abstract (English)
Procedural mistake detection (PMD) is a challenging problem of classifying whether a human user (observed through egocentric video) has successfully executed a task (specified by a procedural text). Despite significant recent efforts, machine performance in the wild remains nonviable, and the reasoning processes underlying this performance are opaque. As such, we extend PMD to require generating visual self-dialog rationales to inform decisions. Given the impressive, mature image understanding capabilities observed in recent vision-and-language models (VLMs), we curate a suitable benchmark dataset for PMD based on individual frames. As our reformulation enables unprecedented transparency, we leverage a natural language inference (NLI) model to formulate two automated metrics for the coherence of generated rationales. We establish baselines for this reframed task, showing that VLMs struggle off-the-shelf, but with some trade-offs, their accuracy, coherence, and efficiency can be improved by incorporating these metrics into common inference and fine-tuning methods. Lastly, our multi-faceted metrics visualize common outcomes, highlighting areas for further improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。