arXiv:2506.22554cs.CVcs.AI2025-06被引 28

构建4000小时真人互动数据集,训练能同步生成动作与表情的对话型虚拟人。

Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset

  • 基于真实对话视频训练双人行为动态模型,支持语音与视觉双向交互
  • 生成动作与表情可随语义和情绪调节,支持2D/3D渲染输出
  • 适用于虚拟助手、远程会议等需要自然互动的场景

人类交流依赖言语与非言语信号的复杂协同,对发展具备社会智能的AI技术至关重要。为此,我们构建了「Seamless Interaction Dataset」,包含超过4,000小时来自4,000名参与者在多样化场景下的面对面互动视频。该数据集支持开发理解双人具身动态的AI技术,推动虚拟代理、远程协作体验和多模态内容分析工具的发展。我们还提出一系列模型,可根据对话者语音与视觉行为生成同步的动作手势与面部表情。模型支持输入来自大语言模型的语音,并集成2D与3D渲染方法,迈向可交互虚拟代理。此外,我们设计了可调控情绪响应与表现力水平的变体模型,提升手势的语义相关性。最后,我们提出了评估此类双人运动模型质量的方法,展示其在实现更直观、响应更灵敏的人机交互中的潜力。

原文摘要 · Abstract (English)

Human communication involves a complex interplay of verbal and nonverbal signals, essential for conveying meaning and achieving interpersonal goals. To develop socially intelligent AI technologies, it is crucial to develop models that can both comprehend and generate dyadic behavioral dynamics. To this end, we introduce the Seamless Interaction Dataset, a large-scale collection of over 4,000 hours of face-to-face interaction footage from over 4,000 participants in diverse contexts. This dataset enables the development of AI technologies that understand dyadic embodied dynamics, unlocking breakthroughs in virtual agents, telepresence experiences, and multimodal content analysis tools. We also develop a suite of models that utilize the dataset to generate dyadic motion gestures and facial expressions aligned with human speech. These models can take as input both the speech and visual behavior of their interlocutors. We present a variant with speech from an LLM model and integrations with 2D and 3D rendering methods, bringing us closer to interactive virtual agents. Additionally, we describe controllable variants of our motion models that can adapt emotional responses and expressivity levels, as well as generating more semantically-relevant gestures. Finally, we discuss methods for assessing the quality of these dyadic motion models, which are demonstrating the potential for more intuitive and responsive human-AI interactions.

虚拟人多模态动作生成交互建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。