arXiv:2604.21291cs.CVcs.AI2026-04

用合成数据提升可控人体视频生成效果,解决真实数据稀缺问题。

Exploring the Role of Synthetic Data Augmentation in Controllable Human-Centric Video Generation

论文配图:Exploring the Role of Synthetic Data Augmentation in Controllable Human-Centric Video Generation
图 1 · 摘自论文原文
  • 基于扩散模型实现外观与动作的精细控制,统一评估合成与真实数据交互
  • 合成数据能显著提升动作真实感、时序一致性与身份保真度
  • 为构建高效通用生成模型提供可落地的数据筛选策略

可控人体视频生成旨在生成具有明确动作和外观引导的逼真人体视频,是数字人、动画和具身AI的基础。然而,大规模、多样且隐私安全的人体视频数据集稀缺,尤其在罕见身份和复杂动作场景下成为主要瓶颈。合成数据提供了可扩展且可控的替代方案,但其对生成建模的实际贡献因持续存在的Sim2Real差距而未被充分探索。本文系统研究了合成数据在可控人体视频生成中的作用。提出一种基于扩散模型的框架,实现外观与动作的细粒度控制,并提供统一测试平台以分析合成数据与真实数据在训练中的交互机制。通过大量实验,揭示了合成与真实数据的互补作用,展示了高效选择合成样本以提升动作真实感、时序一致性和身份保真度的方法。本研究首次全面探索了合成数据在人体中心视频合成中的角色,为构建数据高效且泛化能力强的生成模型提供了实用洞见。

原文摘要 · Abstract (English)

Controllable human video generation aims to produce realistic videos of humans with explicitly guided motions and appearances,serving as a foundation for digital humans, animation, and embodied AI.However, the scarcity of largescale, diverse, and privacy safe human video datasets poses a major bottleneck, especially for rare identities and complex actions.Synthetic data provides a scalable and controllable alternative,yet its actual contribution to generative modeling remains underexplored due to the persistent Sim2Real gap.In this work,we systematically investigate the impact of synthetic data on controllable human video generation. We propose a diffusion-based framework that enables fine-grained control over appearance and motion while providing a unfied testbed to analyze how synthetic data interacts with real world data during training. Through extensive experiments, we reveal the complementary roles of synthetic and real data and demonstrate possible methods for efficiently selecting synthetic samples to enhance motion realism,temporal consistency,and identity preservation.Our study offers the first comprehensive exploration of synthetic data's role in human-centric video synthesis and provides practical insights for building data-efficient and generalizable generative models.

视频生成扩散模型合成数据可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。