用3D轨迹+文本控制生成动态3D形状,让动作更精准贴合描述。
Controllable Dynamic 3D Shape Generation via 3D Trajectories and Text

- 结合3D轨迹与文本,实现空间路径与语义的双重控制。
- 在多个数据集上生成动作更贴合提示,且质量更高。
- 适合需要精确运动控制的动画、游戏和虚拟现实应用。
我们提出T2Mo,一个前馈式框架,通过3D轨迹与文本联合控制动态3D形状生成。由于语言本身存在歧义,仅靠文本生成精确意图的动作仍具挑战。为此,我们引入3D轨迹作为空间引导,明确指定物体上特定点的运动路径。结合两者,T2Mo生成的动作既严格遵循给定轨迹的空间路径,又整体体现文本语义。为应对任意配置的轨迹输入(从稠密到稀疏、分布不均),我们提出一种基于形状的轨迹嵌入方法,将输入轨迹集映射为覆盖整个物体的形状感知标记集。我们在文本基基线和级联视频基基线(结合轨迹引导视频生成与视频转动态网格)上进行广泛对比。定量评估、定性分析及用户研究均表明,本方法生成的动作更忠实于提示,表达力更强,同时保持高质量。
原文摘要 · Abstract (English)
We introduce T2Mo, a feed-forward framework for controllable dynamic 3D shape generation conditioned on 3D trajectories and text. Due to the inherent ambiguity of language, generating precisely intended motions using text alone remains challenging. To address this, we adopt 3D trajectories as controllable spatial guidance, specifying the exact paths along which selected points should move. By combining both, T2Mo generates object motions that spatially adhere to the given trajectories while globally reflecting the text semantics. To robustly handle trajectory inputs with arbitrary configurations, ranging from dense to sparse and unevenly distributed, we further propose a shape-grounded trajectory embedding that maps an input trajectory set into a shape-aware token set covering the entire object. We conduct extensive comparisons against text-based baselines and cascaded video-based baselines that combine trajectory-guided video generation with video-to-dynamic mesh generation. Quantitative and qualitative evaluations, along with user studies, demonstrate that our approach produces motions that more faithfully follow the given prompts with higher expressiveness while preserving motion quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。