arXiv:2511.21998cs.CV2025-11NeurIPS被引 2

评测多模态大模型实时任务指导能力,提出新数据集与实时推理模型。

Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?

  • 构建实时交互式任务指导数据集,含带时间戳的错误提示。
  • 现有模型在实时反馈和纠错上表现不佳,平均延迟超2秒。
  • 适合做智能助手、人机协作系统研发者参考。

多模态大语言模型虽具备较强对话能力,但在提供实时、互动的分步任务指导方面仍存在不足,这要求模型不仅能输出指令,还需实时检测执行效果、识别并预警用户错误,且必须在视频流中异步响应。为此,我们基于CaptainCook4D构建了全新的‘Qualcomm Interactive Cooking’基准与数据集,包含用户执行任务时的错误及纠正过程。该数据集具有密集标注的时间化指令与反馈信息,特别标注了错误在视频中的精确发生时刻。我们在该基准上评估了当前最先进的多模态大模型,并提出LiveMamba——一种专为实时交互式指导设计的流式多模态模型。本工作首次提供了针对实时情境化教学的专用基准与强基线模型。

原文摘要 · Abstract (English)

Multi-modal Large Language Models (LLM) have advanced conversational abilities but struggle with providing live, interactive step-by-step guidance, a key capability for future AI assistants. Effective guidance requires not only delivering instructions but also detecting their successful execution, as well as identifying and alerting users to mistakes, all of which has to happen in real-time. This requires models that are not turn-based, but that can react asynchronously to a video stream, as well as video data showing users performing tasks including mistakes and their corrections. To this end, we introduce Qualcomm Interactive Cooking, a new benchmark and dataset built upon CaptainCook4D, which contains user mistakes during task execution. Our dataset and benchmark features densely annotated, timed instructions and feedback messages, specifically including mistake alerts precisely timestamped to their visual occurrence in the video. We evaluate state-of-the-art multi-modal LLMs on the Qualcomm Interactive Cooking benchmark and introduce LiveMamba, a streaming multi-modal LLM designed for interactive instructional guidance. This work provides the first dedicated benchmark and a strong baseline for developing and evaluating on live, situated coaching.

多模态实时指导任务执行视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。