让智能助手学会在真实场景中判断何时该说话。
EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild
- 从第一视角实时分析视频,决定说话时机。
- 在EasyCom和Ego4D数据集上优于随机和静默基线。
- 适合需要自然对话的户外智能设备使用。
在真实环境中预测何时发起语音交流,仍是对话代理的核心挑战。我们提出EgoSpeak,一种面向第一人称视频流的实时语音发起预测框架。通过从说话者视角建模对话,EgoSpeak适用于需持续观察环境并动态决策的类人交互场景。该方法整合四大能力:第一视角、RGB图像处理、在线处理与未修剪视频处理,有效弥合简化实验与复杂自然对话之间的差距。我们还构建了YT-Conversation——一个来自YouTube的多样化野外对话视频集合,可用于大规模预训练。在EasyCom与Ego4D上的实验表明,EgoSpeak在实时性上显著优于随机与静默基线。结果强调了多模态输入与上下文长度对准确判断说话时机的重要性。
原文摘要 · Abstract (English)
Predicting when to initiate speech in real-world environments remains a fundamental challenge for conversational agents. We introduce EgoSpeak, a novel framework for real-time speech initiation prediction in egocentric streaming video. By modeling the conversation from the speaker's first-person viewpoint, EgoSpeak is tailored for human-like interactions in which a conversational agent must continuously observe its environment and dynamically decide when to talk. Our approach bridges the gap between simplified experimental setups and complex natural conversations by integrating four key capabilities: (1) first-person perspective, (2) RGB processing, (3) online processing, and (4) untrimmed video processing. We also present YT-Conversation, a diverse collection of in-the-wild conversational videos from YouTube, as a resource for large-scale pretraining. Experiments on EasyCom and Ego4D demonstrate that EgoSpeak outperforms random and silence-based baselines in real time. Our results also highlight the importance of multimodal input and context length in effectively deciding when to speak.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。