arXiv:2510.22473cs.CVcs.AI2025-10

从单张图生成高质量4D动态内容,通过姿态对齐提升运动一致性。

DynaPose4D: High-Quality 4D Dynamic Content Generation via Pose Alignment Loss

  • 结合4D高斯点阵与无类别姿态估计,从单图构建3D动态模型。
  • 多视角姿态关键点预测提升运动连贯性,生成动作流畅自然。
  • 适合计算机视觉与动画制作领域,尤其关注动态内容生成的场景。

近年来,2D与3D生成模型推动了计算机视觉的发展。然而,仅凭单张静态图像生成高质量4D动态内容仍是重大挑战。传统方法在建模时间依赖关系和准确捕捉动态几何变化方面存在局限,尤其在相机视角变化时表现不佳。为此,我们提出DynaPose4D,将4D高斯点阵(4DGS)与无类别姿态估计(CAPE)技术相结合。该框架利用3D高斯点阵从单张图像构建3D模型,基于选定视图的一次性支持预测多视角姿态关键点,并引入监督信号增强运动一致性。实验结果表明,DynaPose4D在动态运动生成中表现出优异的连贯性、一致性和流畅性。这些发现不仅验证了该框架的有效性,也展示了其在计算机视觉与动画制作领域的应用潜力。

原文摘要 · Abstract (English)

Recent advancements in 2D and 3D generative models have expanded the capabilities of computer vision. However, generating high-quality 4D dynamic content from a single static image remains a significant challenge. Traditional methods have limitations in modeling temporal dependencies and accurately capturing dynamic geometry changes, especially when considering variations in camera perspective. To address this issue, we propose DynaPose4D, an innovative solution that integrates 4D Gaussian Splatting (4DGS) techniques with Category-Agnostic Pose Estimation (CAPE) technology. This framework uses 3D Gaussian Splatting to construct a 3D model from single images, then predicts multi-view pose keypoints based on one-shot support from a chosen view, leveraging supervisory signals to enhance motion consistency. Experimental results show that DynaPose4D achieves excellent coherence, consistency, and fluidity in dynamic motion generation. These findings not only validate the efficacy of the DynaPose4D framework but also indicate its potential applications in the domains of computer vision and animation production.

4D生成姿态对齐高斯点阵动态建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。