arXiv:2607.02991cs.CV2026-07

首个支持实时任务指导的多领域视频基准,让大模型当实时教练。

GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video

论文配图:GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video
图 1 · 摘自论文原文
  • 构建多领域实时视频交互数据集,支持闭环指导训练与评估。
  • 现有大模型能给指令却难发现错误,纠错能力严重不足。
  • 适合研究实时交互、多模态推理与智能教练系统的学者。

尽管多模态大语言模型(MLLM)在离线视频理解上表现优异,但它们能否充当实时流程指导者仍不明确。这类角色需持续监控执行过程、检测错误并提供纠正指导,形成闭环互动。本文构建了GuideMe,首个支持流式视频的多领域基准,用于训练和评估MLLM在闭环交互任务指导中的能力。数据集包含2,458段视频,覆盖223.7小时,涵盖烹饪、物体操作、日常生活指导和健身等22个领域,共47,775个交互样本,涵盖下一步指令、完成反馈、错误检测与纠正指导。为评估模型,设计三组件评估框架:时序-语义二分匹配(序列对齐)、行为分类(干预时机判断)、大模型作为裁判(内容质量评估)。大量实验揭示关键性能失衡:虽擅长生成指令,现有MLLM普遍无法识别执行错误,难以提供有效纠正反馈。代码与数据已公开于https://fawnliu.github.io/project/guideme。

原文摘要 · Abstract (English)

While multimodal Large Language Models (MLLMs) excel at offline video understanding, an interesting question of how far they are from serving as a real-time procedural coach remains unknown. Such a role typically requires an MLLM to continuously monitor the execution, detect mistakes, and provide corrective guidance in a closed-loop interaction. In this paper, we construct GuideMe, the first multi-domain benchmark for streaming video that supports training and evaluation of MLLMs for closed-loop interactive task guidance. It comprises 2,458 videos spanning 223.7 hours across diverse domains (\eg, cooking, object manipulation, daily-life guidance, and fitness), with 47,775 interaction samples covering next-step instructions, completion feedback, error detection, and corrective guidance. To evaluate existing models on GuideMe, we design a three-component assessment framework to measure the capabilities of representative MLLMs, which consists of temporal-semantic bipartite matching for sequence-level alignment, behavioral classification for intervention timing, and LLM-as-a-Judge for content quality. Extensive experiments highlight a critical performance asymmetry: despite excelling at providing instructions, existing MLLMs consistently fail to identify execution errors and respond with corrective feedback. Code and data are released at https://fawnliu.github.io/project/guideme.

视频生成多模态智能教练闭环交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。