arXiv:2507.20987cs.CVcs.AI2025-07

首个联合生成全身动作与语音的基准数据集,解决多模态一致性难题。

JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1

  • 构建10,000身份、200万视频的大规模多模态数据集
  • 发现面部/手部主导模型在全身生成上性能显著下降
  • 公开数据集与评测工具,适合音视频生成研究者使用

基于扩散模型的视频生成虽已实现高保真短片段,但联合生成全身动作与自然语音时仍难以保证多模态一致性。现有方法缺乏全面评估框架,且缺少针对特定区域的性能分析基准。为此,我们推出首个联合全身说话虚拟人与语音生成基准JWB-DH-V1,包含10,000个唯一身份、200万视频样本的多模态数据集,以及评估全身可驱动虚拟人音视频联合生成的评测协议。对当前最优模型的评估显示,面部/手部中心模型在全身生成任务中表现明显落后,揭示了未来研究的关键方向。数据集与评测工具已在https://github.com/deepreasonings/WholeBodyBenchmark公开。

原文摘要 · Abstract (English)

Recent advances in diffusion-based video generation have enabled photo-realistic short clips, but current methods still struggle to achieve multi-modal consistency when jointly generating whole-body motion and natural speech. Current approaches lack comprehensive evaluation frameworks that assess both visual and audio quality, and there are insufficient benchmarks for region-specific performance analysis. To address these gaps, we introduce the Joint Whole-Body Talking Avatar and Speech Generation Version I(JWB-DH-V1), comprising a large-scale multi-modal dataset with 10,000 unique identities across 2 million video samples, and an evaluation protocol for assessing joint audio-video generation of whole-body animatable avatars. Our evaluation of SOTA models reveals consistent performance disparities between face/hand-centric and whole-body performance, which incidates essential areas for future research. The dataset and evaluation tools are publicly available at https://github.com/deepreasonings/WholeBodyBenchmark.

虚拟人生成多模态生成音视频同步

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。