零样本文本生成动态网格,解决形状一致性与几何保真难题
TextMesh4D: Zero-shot Text-to-4D Mesh Generation
- 用面片变形场替代顶点变形,保持拓扑结构稳定
- 在单张24GB GPU上实现高质量时序一致的4D网格生成
- 适合需要高保真动态3D模型的虚拟人、动画生成场景
大规模高质量动态3D(4D)资产对学习物理基础表征至关重要,但其采集和标注成本高昂,限制了监督式4D学习。为此,本文提出零样本文本到4D网格生成方法。现有方法多采用隐式3D表示(如NeRF或3DGS),虽具强形变能力,但难以控制表面拓扑,导致几何保真度低且时序重建困难。本文提出TextMesh4D,首个直接生成动态网格的零样本框架。针对扩散引导与拓扑约束间的不匹配问题,从几何与语义双维度改进:几何上,将形变建模从顶点转向面片,通过雅可比变形场(JDF)结合可积性约束实现拓扑感知重建;语义上,引入局部-全局语义正则化器(LGSR),联合约束局部形变合理性与全局形状一致性以保持身份连续性。大量实验表明,该方法在时序一致性、结构保真度与视觉质量上均达当前最优,且仅需单张24GB GPU即可高效运行。
原文摘要 · Abstract (English)
Large-scale, high-quality dynamic 3D (4D) assets are essential for learning physically grounded representations, but remain costly to capture and annotate at scale. This limits the viability of supervised 4D learning and motivates zero-shot text-to-4D generation leveraging pretrained diffusion priors. To model complex dynamics, prior methods typically adopt implicit 3D representations (e.g., NeRFs or 3DGS) for their deformation capacity. However, their implicit nature provides limited control over surface topology, which hinders high-fidelity geometry and makes temporally coherent surface reconstruction challenging. To address these limitations, we explore zero-shot text-to-4D mesh generation. However, a structural mismatch arises when combining diffusion-based guidance with topology-constrained meshes: the guidance is noisy and spatially inconsistent, while meshes impose severe topological constraints, making direct vertex-level deformation unstable. In this paper, we introduce TextMesh4D, the first zero-shot framework for text-to-4D that directly generates dynamic meshes by addressing the above challenge at two complementary levels. Geometrically, we shift deformation modeling from vertices to faces via a Jacobian Deformation Field (JDF), enabling topology-aware surface reconstruction through an integrability-enforcing integration formulation. Semantically, we propose a Local-Global Semantic Regularizer (LGSR) that preserves identity over time by jointly constraining local deformation plausibility and global shape consistency. Extensive experiments demonstrate state-of-the-art temporal consistency, structural fidelity, and visual quality, while remaining efficient on a single 24GB GPU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。