arXiv:2506.05397cs.GRcs.AI2025-06被引 4

自动生成逼真4D人体动画,解决野外场景数据稀缺问题

Gen4D: Synthesizing Humans and Scenes in the Wild

  • 用扩散模型生成多样人体形象,结合动作编码与环境合成
  • 构建涵盖棒球、冰球、足球的大型合成数据集SportPAL
  • 无需手动建模,适合研究野外人类行为的视觉任务

野外活动的数据匮乏导致计算机视觉任务性能受限,尤其在体育等非常见人类主导领域,真实数据采集复杂且不现实。现有合成方法因依赖固定资产库和手工渲染流程,普遍存在人体外观、动作和场景组合多样性不足的问题。为此,我们提出Gen4D,一个全自动生成多样化、逼真4D人体动画的流水线。该方法融合专家驱动的动作编码、基于扩散的高斯点阵提示引导的人体生成,以及人体感知的背景合成技术,生成高度多变且逼真的动态序列。基于Gen4D,我们构建了SportPAL——一个覆盖棒球、冰球和足球三个运动的大规模合成数据集。Gen4D与SportPAL共同为野外人类主导的视觉任务提供可扩展的合成数据基础,无需人工3D建模或场景设计。

原文摘要 · Abstract (English)

Lack of input data for in-the-wild activities often results in low performance across various computer vision tasks. This challenge is particularly pronounced in uncommon human-centric domains like sports, where real-world data collection is complex and impractical. While synthetic datasets offer a promising alternative, existing approaches typically suffer from limited diversity in human appearance, motion, and scene composition due to their reliance on rigid asset libraries and hand-crafted rendering pipelines. To address this, we introduce Gen4D, a fully automated pipeline for generating diverse and photorealistic 4D human animations. Gen4D integrates expert-driven motion encoding, prompt-guided avatar generation using diffusion-based Gaussian splatting, and human-aware background synthesis to produce highly varied and lifelike human sequences. Based on Gen4D, we present SportPAL, a large-scale synthetic dataset spanning three sports: baseball, icehockey, and soccer. Together, Gen4D and SportPAL provide a scalable foundation for constructing synthetic datasets tailored to in-the-wild human-centric vision tasks, with no need for manual 3D modeling or scene design.

4D生成合成数据人体动画扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。