arXiv:2504.18318cs.CV2025-04被引 5

让文字生成的4D内容更连贯真实,解决动态画面不一致问题。

STP4D: Spatio-Temporal-Prompt Consistent Modeling for Text-to-4D Gaussian Splatting

  • 设计时序提示嵌入与几何增强模块,实现时空与文本的一致性建模。
  • 生成1个4D资产仅需约4.6秒,速度远超现有方法。
  • 首次用扩散模型生成4D高斯点,兼顾质量与实时渲染效率。

文本到4D生成技术快速发展并广泛应用于多个场景。然而,现有方法往往缺乏统一框架下的时空建模与提示对齐,导致时间不连贯、几何失真或生成内容偏离文本描述。为此,我们提出STP4D,一种集成全面时空-提示一致性建模的新方法。该方法包含三个精心设计模块:时变提示嵌入、几何信息增强和时序扩展形变,协同实现高质量4D生成。此外,STP4D是首批将扩散模型用于生成4D高斯点的方法之一,结合4DGS的精细建模能力与实时渲染优势,以及扩散模型的快速推理速度。大量实验表明,STP4D在生成高保真4D内容方面表现卓越,效率极高(每资产约4.6秒),在质量和速度上均优于现有方法。

原文摘要 · Abstract (English)

Text-to-4D generation is rapidly developing and widely applied in various scenarios. However, existing methods often fail to incorporate adequate spatio-temporal modeling and prompt alignment within a unified framework, resulting in temporal inconsistencies, geometric distortions, or low-quality 4D content that deviates from the provided texts. Therefore, we propose STP4D, a novel approach that aims to integrate comprehensive spatio-temporal-prompt consistency modeling for high-quality text-to-4D generation. Specifically, STP4D employs three carefully designed modules: Time-varying Prompt Embedding, Geometric Information Enhancement, and Temporal Extension Deformation, which collaborate to accomplish this goal. Furthermore, STP4D is among the first methods to exploit the Diffusion model to generate 4D Gaussians, combining the fine-grained modeling capabilities and the real-time rendering process of 4DGS with the rapid inference speed of the Diffusion model. Extensive experiments demonstrate that STP4D excels in generating high-fidelity 4D content with exceptional efficiency (approximately 4.6s per asset), surpassing existing methods in both quality and speed.

4D生成扩散模型高斯溅射文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。