arXiv:2602.01038cs.CV2026-02

将单人教学视频自动生成双人对话,助力任务辅助AI训练

From Videos to Conversations: Egocentric Instructions for Task Assistance

  • 用大模型自动把单人教学视频转成多人多模态对话
  • 构建了含507段对话、6636个问答的跨领域数据集
  • 适合做任务指导类AI研究者参考

日常任务如家电维修、烹饪和汽车保养等,常需专家知识,尤其涉及复杂多步骤流程。尽管人工智能代理在增强现实(AR)辅助领域备受关注,但进展受限于真实世界任务执行中大规模多模态对话数据集的缺乏,主要因人工数据采集成本高、流程复杂。本文提出一种全自动框架,可将单人教学视频转化为双人多模态任务指导对话。该框架基于大语言模型,实现可扩展且低成本的数据生成。基于此,我们构建了HowToDIV数据集,包含507段对话、6,636个问答对及24小时视频,覆盖多个领域。每段对话为多轮专家-新手交互。最后,我们在HowToDIV上使用Gemma 3和Qwen 2.5报告基线结果,为多模态程序性任务辅助提供初步基准。

原文摘要 · Abstract (English)

Many everyday tasks, ranging from appliance repair and cooking to car maintenance, require expert knowledge, particularly for complex, multi-step procedures. Despite growing interest in AI agents for augmented reality (AR) assistance, progress remains limited by the scarcity of large-scale multimodal conversational datasets grounded in real-world task execution, in part due to the cost and logistical complexity of human-assisted data collection. In this paper, we present a framework to automatically transform single person instructional videos into two-person multimodal task-guidance conversations. Our fully automatic pipeline, based on large language models, provides a scalable and cost efficient alternative to traditional data collection approaches. Using this framework, we introduce HowToDIV, a multimodal dataset comprising 507 conversations, 6,636 question answer pairs, and 24 hours of video spanning multiple domains. Each session consists of a multi-turn expert-novice interaction. Finally, we report baseline results using Gemma 3 and Qwen 2.5 on HowToDIV, providing an initial benchmark for multimodal procedural task assistance.

任务辅助多模态对话数据生成视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。