arXiv:2510.25319cs.GRcs.AI2025-10

用文字生成会动的3D手绘草图,轻量易懂适合设计原型。

4-Doodle: Text to 3D Sketches that Move!

  • 用双空间蒸馏融合图像视频模型,通过贝塞尔曲线建模多视角一致的3D结构。
  • 无需训练,生成动画在时间上连贯、结构稳定,支持翻转旋转等复杂动作。
  • 适合交互设计、创意草图快速可视化,不依赖专业建模工具。

我们提出一项新任务:文本到动态3D手绘草图生成,旨在将自由手绘草图在动态3D空间中“活化”。与以往聚焦于照片级真实感内容生成的工作不同,我们关注稀疏、风格化且多视角一致的3D矢量草图,这种轻量化、可解释的媒介非常适合视觉表达与原型设计。然而该任务极具挑战:(i) 不存在文本与3D(或4D)草图的配对数据集;(ii) 草图需结构抽象,传统3D表示如NeRF或点云难以建模;(iii) 动画需保证时间连贯性和多视角一致性,现有流程无法解决。为此,我们提出4-Doodle,首个无训练的文本到动态3D草图生成框架。它通过双空间蒸馏方案利用预训练图像和视频扩散模型:一空间使用可微贝塞尔曲线捕捉多视角一致几何,另一空间通过时序感知先验编码运动动态。不同于先前方法(如DreamFusion)每步仅优化单视角,我们的多视角优化确保结构对齐并避免视图歧义,这对稀疏草图至关重要。此外,我们引入结构感知运动模块,分离形状保持轨迹与形变相关变化,实现翻转、旋转及关节运动等丰富动作。大量实验表明,本方法生成的动画在时间真实性与结构稳定性上均优于现有基线,在保真度与可控性上表现更优。我们希望此工作能推动更直观、易用的4D内容创作发展。

原文摘要 · Abstract (English)

We present a novel task: text-to-3D sketch animation, which aims to bring freeform sketches to life in dynamic 3D space. Unlike prior works focused on photorealistic content generation, we target sparse, stylized, and view-consistent 3D vector sketches, a lightweight and interpretable medium well-suited for visual communication and prototyping. However, this task is very challenging: (i) no paired dataset exists for text and 3D (or 4D) sketches; (ii) sketches require structural abstraction that is difficult to model with conventional 3D representations like NeRFs or point clouds; and (iii) animating such sketches demands temporal coherence and multi-view consistency, which current pipelines do not address. Therefore, we propose 4-Doodle, the first training-free framework for generating dynamic 3D sketches from text. It leverages pretrained image and video diffusion models through a dual-space distillation scheme: one space captures multi-view-consistent geometry using differentiable Bézier curves, while the other encodes motion dynamics via temporally-aware priors. Unlike prior work (e.g., DreamFusion), which optimizes from a single view per step, our multi-view optimization ensures structural alignment and avoids view ambiguity, critical for sparse sketches. Furthermore, we introduce a structure-aware motion module that separates shape-preserving trajectories from deformation-aware changes, enabling expressive motion such as flipping, rotation, and articulated movement. Extensive experiments show that our method produces temporally realistic and structurally stable 3D sketch animations, outperforming existing baselines in both fidelity and controllability. We hope this work serves as a step toward more intuitive and accessible 4D content creation.

3D草图文本生成动态图形扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。