将单人教学视频自动转为任务指导对话,构建大规模对话-视频数据集。
Generating Dialogues from Egocentric Instructional Videos for Task Assistance: Dataset, Method and Benchmark
- 用大模型自动将单人视频转为专家-新手对话,无需人工标注
- 构建包含507段对话、6636个问答对的HowToDIV数据集,覆盖24小时视频
- 适合研究任务型对话生成、AI助教系统的人群使用
日常任务如修电器、做菜、汽车保养等常需专业知识,尤其在复杂多步骤场景下。尽管人工智能代理研究兴起,但面向真实世界任务协助的对话-视频数据集仍十分稀缺。本文提出一种简单有效的自动化方法,将单人教学视频转换为与细粒度步骤和视频片段对齐的双人任务指导对话。该方法基于大语言模型,完全自动,显著降低人力成本。基于此,我们构建了HowToDIV数据集,涵盖507段对话、6636个问答对和24小时视频片段,覆盖烹饪、机械、种植等多样化任务。每段会话包含多轮对话,专家通过可穿戴设备的摄像头和麦克风观察用户环境,逐步指导其完成任务。我们以Gemma-3模型建立基准,为未来任务型对话生成研究提供新方向。
原文摘要 · Abstract (English)
Many everyday tasks ranging from fixing appliances, cooking recipes to car maintenance require expert knowledge, especially when tasks are complex and multi-step. Despite growing interest in AI agents, there is a scarcity of dialogue-video datasets grounded for real world task assistance. In this paper, we propose a simple yet effective approach that transforms single-person instructional videos into task-guidance two-person dialogues, aligned with fine grained steps and video-clips. Our fully automatic approach, powered by large language models, offers an efficient alternative to the substantial cost and effort required for human-assisted data collection. Using this technique, we build HowToDIV, a large-scale dataset containing 507 conversations, 6636 question-answer pairs and 24 hours of videoclips across diverse tasks in cooking, mechanics, and planting. Each session includes multi-turn conversation where an expert teaches a novice user how to perform a task step by step, while observing user's surrounding through a camera and microphone equipped wearable device. We establish the baseline benchmark performance on HowToDIV dataset through Gemma-3 model for future research on this new task of dialogues for procedural-task assistance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。