arXiv:2507.09862cs.CVeess.AS2025-07被引 25

首个大规模音视频双人交互虚拟人数据集,支持高保真数字人对话生成。

SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation

  • 构建520万段音视频片段,覆盖对话、独白、倾听等四种互动类型。
  • 分设大规模预训练与高质量精调子集,支持不同阶段模型训练。
  • 配套基准测试和自回归视频聊天基线,助力虚拟人交互研究。

大规模模型的快速发展推动了数字人领域的重大突破。这些先进方法为虚拟人驱动与渲染提供了高保真解决方案,促使学界聚焦于下一代核心挑战:音视频双人交互式虚拟人生成。为此,我们提出SpeakerVid-5M数据集,这是首个专为音视频双人交互式虚拟人生成设计的大规模、高质量数据集。该数据集总时长超过8,743小时,包含超过520万段人物视频片段,涵盖单人说话、倾听及双人对话等多种互动场景。数据按两个关键维度组织:互动类型(对话分支、单人分支、倾听分支、多轮分支)与数据质量(大规模预训练子集与精选高质量子集)。这一双重结构可适配多种2D虚拟人任务。此外,我们基于该数据集训练了一个自回归视频聊天基线模型,并提供专用评估指标与测试数据集,作为未来研究的基准 VidChatBench。数据集及处理代码将公开发布。

原文摘要 · Abstract (English)

The rapid development of large-scale models has catalyzed significant breakthroughs in the digital human domain. These advanced methodologies offer high-fidelity solutions for avatar driving and rendering, leading academia to focus on the next major challenge: audio-visual dyadic interactive virtual human. To facilitate research in this emerging area, we present SpeakerVid-5M dataset, the first large-scale, high-quality dataset designed for audio-visual dyadic interactive virtual human generation. Totaling over 8,743 hours, SpeakerVid-5M contains more than 5.2 million video clips of human portraits. It covers diverse scales and interaction types, including monadic talking, listening, and dyadic conversations. Crucially, the dataset is structured along two key dimensions: interaction type and data quality. First, it is categorized into four types (dialogue branch, single branch, listening branch and multi-turn branch) based on the interaction scenario. Second, it is stratified into a large-scale pre-training subset and a curated, high-quality subset for Supervised Fine-Tuning (SFT). This dual structure accommodates a wide array of 2D virtual human tasks. In addition, we provide an autoregressive (AR)-based video chat baseline trained on this data, accompanied by a dedicated set of metrics and test data to serve as a benchmark VidChatBench for future work. Both the dataset and the corresponding data processing code will be publicly released. Project page: https://dorniwang.github.io/SpeakerVid-5M/

虚拟人音视频数据集对话生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。