arXiv:2412.20422cs.CV2024-12被引 4

无需训练,用文本控制3D物体动态生成,保持原形且动作自然。

Bringing Objects to Life: training-free 4D generation from 3D objects through view consistent noise

  • 将3D网格转为静态4D NeRF,再用文生视频扩散模型驱动动画
  • 通过视角一致的噪声策略提升运动真实感,视觉保真度更高
  • 适合需要快速生成定制化动态3D内容的设计师与开发者

近期生成模型进展使得基于文本提示生成动态4D内容(即运动中的3D物体)成为可能,广泛应用于虚拟世界、媒体与游戏。现有方法虽能控制生成内容外观并动画化3D物体,但其动态能力受限于训练时使用的网格数据集,缺乏生长或结构演变能力。本文提出一种无需训练的方法:通过文本提示引导4D生成,实现自定义通用场景的同时保持原始物体身份。首先将3D网格转换为保留视觉属性的静态4D Neural Radiance Field(NeRF),随后利用图像到视频扩散模型进行动画生成。为提升运动真实性,引入视角一致的噪声协议,使物体视角与噪声过程对齐以促进自然运动;并采用掩码得分蒸馏采样(masked Score Distillation Sampling, SDS)损失,借助注意力图聚焦优化关键区域,更好保留原始物体特征。我们在两个不同3D物体数据集上评估了时间一致性、提示遵循性与视觉保真度,结果表明,相比基于多视图训练的基线方法,本方法在困难场景下更契合文本提示,表现更优。

原文摘要 · Abstract (English)

Recent advancements in generative models have enabled the creation of dynamic 4D content - 3D objects in motion - based on text prompts, which holds potential for applications in virtual worlds, media, and gaming. Existing methods provide control over the appearance of generated content, including the ability to animate 3D objects. However, their ability to generate dynamics is limited to the mesh datasets they were trained on, lacking any growth or structural development capability. In this work, we introduce a training-free method for animating 3D objects by conditioning on textual prompts to guide 4D generation, enabling custom general scenes while maintaining the original object's identity. We first convert a 3D mesh into a static 4D Neural Radiance Field (NeRF) that preserves the object's visual attributes. Then, we animate the object using an Image-to-Video diffusion model driven by text. To improve motion realism, we introduce a view-consistent noising protocol that aligns object perspectives with the noising process to promote lifelike movement, and a masked Score Distillation Sampling (SDS) loss that leverages attention maps to focus optimization on relevant regions, better preserving the original object. We evaluate our model on two different 3D object datasets for temporal coherence, prompt adherence, and visual fidelity, and find that our method outperforms the baseline based on multiview training, achieving better consistency with the textual prompt in hard scenarios.

4D生成扩散模型零样本神经辐射场

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。