arXiv:2607.27634cs.CV2026-07

直接从文本生成360度动态人类,1分钟完成且视图一致。

4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans

论文配图:4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans
图 1 · 摘自论文原文
  • 端到端建模4D高斯点云,跳过视频预生成和重建
  • 1分钟生成,时间效率提升10倍以上,多视角一致性更好
  • 适合影视、游戏等需要高质量动态角色的场景

从文本提示生成高质量360度动态人类资产极具挑战性。现有方法通常先合成单目或多视角视频,再拟合4D表示,过程昂贵且易导致几何不完整或视角不一致。我们提出4DHumanDiff,一种直接从文本生成4D高斯点云(4DGS)表示的扩散框架。通过端到端建模结构化的4D表示空间,4DHumanDiff避免了视频预生成和场景级重建,更适用于视图一致且时序连贯的资产生成。模型采用带时间注意力的3D U-Net骨干网络实现运动感知生成。我们构建了一个包含6万对高质量数据的文本到4DGS数据集,并引入2D正则化和无需训练的4D插值方法以提升渲染质量和运动平滑性。实验表明,4DHumanDiff可在1分钟内生成一致的360度动态人类,实现更好的时空一致性,推理时间减少超过10倍。

原文摘要 · Abstract (English)

Generating high-quality 360-degree dynamic human assets from text prompts is challenging. Existing methods usually synthesize monocular or multi-view videos first and then fit a 4D representation, which is expensive and often causes incomplete geometry or view-inconsistent renderings. We present 4DHumanDiff, a diffusion framework that directly generates dynamic humans represented by 4D Gaussian Splatting (4DGS) from text prompts. By modeling the structured 4D representation space end-to-end, 4DHumanDiff avoids video pre-generation and per-scene reconstruction, making it better suited for view-consistent and temporally coherent asset generation. The model uses a 3D U-Net backbone with temporal attention for motion-aware generation. We further construct a large-scale text-to-4DGS dataset with 60,000 high-quality pairs, and introduce 2D regularization and training-free 4D interpolation to improve rendering quality and motion smoothness. Experiments show that 4DHumanDiff generates consistent 360-degree dynamic humans within one minute, achieves better temporal and multi-view consistency, and reduces inference time by more than 10x.

4D高斯文本生成动态人体扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。