arXiv:2509.21888cs.CV2025-09被引 1

让用户用文字控制3D物体运动,精准融入高质量3D场景。

Drag4D: Align Your Motion with Text-Driven 3D Scene Generation

论文配图:Drag4D: Align Your Motion with Text-Driven 3D Scene Generation
图 1 · 摘自论文原文
  • 通过3D复制粘贴和物理感知定位,实现物体在3D场景中的精确空间对齐。
  • 采用多视角视频扩散模型,支持用户定义的3D轨迹并保持运动一致性。
  • 适用于需要交互式3D内容创作的设计师与开发者,尤其适合影视动画场景生成。

我们提出Drag4D,一个将物体运动控制融入文本驱动3D场景生成的交互式框架。该框架允许用户基于单张图像定义3D物体的运动轨迹,并将其无缝整合至高质量3D背景中。整个流程分为三阶段:第一阶段利用全景图与修复的新视角,结合2D高斯泼溅增强文本到3D背景生成,实现稠密且视觉完整的3D重建;第二阶段,基于目标物体参考图像,采用现成的图像到3D模型提取完整3D网格,并通过物理感知位置学习实现与3D场景的精准空间对齐;第三阶段,根据用户定义的3D轨迹对已对齐物体进行时序动画。为缓解运动幻觉并确保视图一致的时序对齐,我们设计了基于部分增强与运动条件的视频扩散模型,联合处理多视角图像对及其投影2D轨迹。我们在各阶段及最终结果上均进行评估,验证了统一架构在用户可控物体运动与高质量3D背景协同对齐方面的有效性。

原文摘要 · Abstract (English)

We introduce Drag4D, an interactive framework that integrates object motion control within text-driven 3D scene generation. This framework enables users to define 3D trajectories for the 3D objects generated from a single image, seamlessly integrating them into a high-quality 3D background. Our Drag4D pipeline consists of three stages. First, we enhance text-to-3D background generation by applying 2D Gaussian Splatting with panoramic images and inpainted novel views, resulting in dense and visually complete 3D reconstructions. In the second stage, given a reference image of the target object, we introduce a 3D copy-and-paste approach: the target instance is extracted in a full 3D mesh using an off-the-shelf image-to-3D model and seamlessly composited into the generated 3D scene. The object mesh is then positioned within the 3D scene via our physics-aware object position learning, ensuring precise spatial alignment. Lastly, the spatially aligned object is temporally animated along a user-defined 3D trajectory. To mitigate motion hallucination and ensure view-consistent temporal alignment, we develop a part-augmented, motion-conditioned video diffusion model that processes multiview image pairs together with their projected 2D trajectories. We demonstrate the effectiveness of our unified architecture through evaluations at each stage and in the final results, showcasing the harmonized alignment of user-controlled object motion within a high-quality 3D background.

3D生成运动控制视频扩散交互生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。