arXiv:2508.03050cs.CV2025-08被引 1

构建首个大规模多人对话视频数据集,推动真实互动场景生成研究

Multi-human Interactive Talking Dataset

  • 自动采集并标注12小时多人对话视频,支持2-4人交互
  • 提出CovOG模型,融合姿态编码与语音驱动实现动态头部动作
  • 适用于多说话人视频生成、人机交互等前沿方向研究

现有对话视频生成研究多集中于单人独白或孤立面部动画,难以应对真实多人交互场景。为此,我们提出MIT——一个专为多人对话视频生成设计的大规模数据集。通过自动化流程收集并标注多人群体对话视频,数据集包含12小时高清视频,每段视频含2至4名说话人,并提供精细的肢体姿态与语音交互标注,捕捉真实对话中的多说话人动态行为,为研究交互视觉行为提供丰富资源。为验证MIT潜力,我们进一步提出CovOG基准模型,其包含多人体姿态编码器(MPE)以处理不同数量说话人,通过聚合个体姿态嵌入实现统一建模;以及交互式音频驱动模块(IAD),基于说话人特异性音频特征调控头部运动。二者协同展示了生成逼真多人对话视频的可行性与挑战,确立了MIT作为未来研究的重要基准。代码已开源:https://github.com/showlab/Multi-human-Talking-Video-Dataset。

原文摘要 · Abstract (English)

Existing studies on talking video generation have predominantly focused on single-person monologues or isolated facial animations, limiting their applicability to realistic multi-human interactions. To bridge this gap, we introduce MIT, a large-scale dataset specifically designed for multi-human talking video generation. To this end, we develop an automatic pipeline that collects and annotates multi-person conversational videos. The resulting dataset comprises 12 hours of high-resolution footage, each featuring two to four speakers, with fine-grained annotations of body poses and speech interactions. It captures natural conversational dynamics in multi-speaker scenario, offering a rich resource for studying interactive visual behaviors. To demonstrate the potential of MIT, we furthur propose CovOG, a baseline model for this novel task. It integrates a Multi-Human Pose Encoder (MPE) to handle varying numbers of speakers by aggregating individual pose embeddings, and an Interactive Audio Driver (IAD) to modulate head dynamics based on speaker-specific audio features. Together, these components showcase the feasibility and challenges of generating realistic multi-human talking videos, establishing MIT as a valuable benchmark for future research. The code is avalibale at: https://github.com/showlab/Multi-human-Talking-Video-Dataset.

对话视频多人交互数据集姿态建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。