用第一视角视频生成实时主动对话助手,解决数据难收集问题
Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
- 从标注的第一视角视频合成大规模对话数据
- 提出自动评估指标并通过真人测试验证有效性
- 端到端模型支持长视频流,能主动指导用户完成任务
近年来对话式AI进展显著,但实现实时感知任务引导系统仍具挑战。这类系统需基于持续输入的视觉信息提供交互式、主动式协助,但其发展受限于数据收集与系统评估成本高昂。为此,本文提出一个完整框架,包含三大贡献:首先,构建新颖的数据清洗流程,从标注的第一人称视频中合成对话,形成覆盖多领域的大型合成对话数据集 extit{dataset};其次,开发一套经大量人工实验验证的自动评估指标;第三,提出一个端到端模型,可处理流式视频输入,生成上下文相关的响应,并引入新方法应对数据不平衡与长时视频问题。本工作为构建实时、主动型AI助手奠定了基础。
原文摘要 · Abstract (English)
Recent advances in conversational AI have been substantial, but developing real-time systems for perceptual task guidance remains challenging. These systems must provide interactive, proactive assistance based on streaming visual inputs, yet their development is constrained by the costly and labor-intensive process of data collection and system evaluation. To address these limitations, we present a comprehensive framework with three key contributions. First, we introduce a novel data curation pipeline that synthesizes dialogues from annotated egocentric videos, resulting in \dataset, a large-scale synthetic dialogue dataset spanning multiple domains. Second, we develop a suite of automatic evaluation metrics, validated through extensive human studies. Third, we propose an end-to-end model that processes streaming video inputs to generate contextually appropriate responses, incorporating novel techniques for handling data imbalance and long-duration videos. This work lays the foundation for developing real-time, proactive AI assistants capable of guiding users through diverse tasks. Project page: https://pro-assist.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。