arXiv:2606.09547cs.CVcs.LG2026-06

视频大模型能实时纠错,帮人学做菜。

Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur?

论文配图:Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur?
图 1 · 摘自论文原文
  • 构建真实烹饪场景下的错误纠正基准,评估模型实时干预能力。
  • 现有视频大模型在该任务上表现差,因缺乏带纠错标注的数据。
  • 生成合成数据集Ego-CoMist,提升小模型在边缘设备上的纠错效果。

学习日常技能如烹饪,越来越多依赖在线视频等教学媒体。这为视频(及多模态)大语言模型作为任务指导助手提供了可能。一个关键能力是:一旦发现用户操作出错,便能及时主动干预。为此,我们提出Ego-MC-Bench(错误纠正)基准,用于评估真实烹饪场景中反应式、分步式任务指导的性能。大量实验表明,该基准对当前最先进的视频大模型极具挑战性。我们认为主要原因在于训练数据有限——尽管存在大量烹饪视频数据集,但缺乏包含错误及恰当干预时机的标注样本。为解决这一问题,我们引入Ego-CoMist,一个通过转化非交互式烹饪视频生成的反事实合成数据集,用于监督训练模型实现主动干预。实验显示,在Ego-CoMist上微调可显著提升较小、更高效的视频大模型性能,使其更适合在边缘设备上部署。

原文摘要 · Abstract (English)

Learning everyday skills, like cooking a dish, relies increasingly on instructional media such as online videos. This opens the door to the use of video (and multimodal) large language models (LLMs) as task guidance assistants. A crucial capability for the real-world success of a prospective task guidance assistant is it's ability to intervene proactively as soon as a mistake is apparent in order to guide the user. To evaluate this crucial capability, we introduce Ego-MC-Bench (Mistake Corrections), a benchmark for evaluating reactive, step-by-step task guidance in realistic cooking scenarios. Extensive experiments show that Ego-MC-Bench is highly challenging for state-of-the-art video LLMs. We argue that a key reason is the limited availability of training data for fine-tuning models on this task. Although there exists a wide range of cooking video datasets, existing datasets lack examples of mistakes along with appropriately timed interventions. To help address this data limitation, we also introduce Ego-CoMist, a counterfactual synthetic dataset created by transforming non -interactive cooking videos into supervised training examples showing proactive interventions. We show that fine-tuning on Ego-CoMist yields performance gains especially for smaller and more efficient video LLMs that are well suited for delivering assistance on edge devices.

视频大模型实时纠错边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。