arXiv:2602.19146cs.CVcs.CL2026-02Conference of the …

让AI能看视频、懂步骤、会对话,精准指导做饭修物

VIGiA: Instructional Video Guidance via Dialogue Reasoning and Retrieval

  • 结合视觉与语言理解,实现对复杂任务计划的多步推理
  • 在任务问答中达到90%以上准确率,显著优于现有模型
  • 适合需要交互式操作指导的应用,如教学或智能助手

我们提出VIGiA,一种新型多模态对话模型,用于理解和推理复杂的多步骤指令视频动作规划。与以往主要关注文本引导或孤立处理视觉与语言的工作不同,VIGiA支持基于场景的、计划感知的对话,需对视觉输入、任务计划及用户交互进行综合推理。为此,VIGiA引入两项关键能力:(1)多模态计划推理,使模型能够将单模态或跨模态查询与当前任务计划对齐并准确回应;(2)基于计划的检索,可从文本或视觉表示中检索相关计划步骤。我们在一个新构建的数据集上进行了实验,该数据集包含与烹饪和DIY计划对齐的丰富指令视频对话。评估结果表明,VIGiA在所有对话式计划引导任务中均优于现有最先进模型,在计划感知视觉问答任务中准确率超过90%。

原文摘要 · Abstract (English)

We introduce VIGiA, a novel multimodal dialogue model designed to understand and reason over complex, multi-step instructional video action plans. Unlike prior work which focuses mainly on text-only guidance, or treats vision and language in isolation, VIGiA supports grounded, plan-aware dialogue that requires reasoning over visual inputs, instructional plans, and interleaved user interactions. To this end, VIGiA incorporates two key capabilities: (1) multimodal plan reasoning, enabling the model to align uni- and multimodal queries with the current task plan and respond accurately; and (2) plan-based retrieval, allowing it to retrieve relevant plan steps in either textual or visual representations. Experiments were done on a novel dataset with rich Instructional Video Dialogues aligned with Cooking and DIY plans. Our evaluation shows that VIGiA outperforms existing state-of-the-art models on all tasks in a conversational plan guidance setting, reaching over 90\% accuracy on plan-aware VQA.

视频理解多模态对话任务规划智能指导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。