arXiv:2506.22432cs.CV2025-06SIGGRAPH被引 12

用3D模型精准控制视频编辑,保持动作一致性。

Shape-for-Motion: Precise and Consistent Video Editing with 3D Proxy

  • 通过3D网格代理实现帧间一致的视频编辑
  • 单帧编辑可自动传播至全视频,减少用户操作
  • 支持姿态、纹理、组合等物理合理修改,适合影视创作

深度生成建模的进步为视频合成带来了前所未有的机遇。但在实际应用中,用户常需精确且一致地实现创作意图。现有方法虽有进展,但确保与用户意图的细粒度对齐仍是难题。本文提出Shape-for-Motion,通过将输入视频中的目标物体转换为时序一致的3D网格(即3D代理),使编辑操作直接在代理上进行,再反推回视频帧。我们设计了双传播策略,用户只需在单帧3D网格上编辑,即可自动传播至其他帧。不同帧的3D网格投影到2D空间生成编辑后的几何与纹理渲染图,作为解耦视频扩散模型的输入,生成最终结果。该框架支持多种精确且物理一致的操作,包括姿态调整、旋转、缩放、平移、纹理修改及物体组合。大量实验验证了方法的有效性与优越性。

原文摘要 · Abstract (English)

Recent advances in deep generative modeling have unlocked unprecedented opportunities for video synthesis. In real-world applications, however, users often seek tools to faithfully realize their creative editing intentions with precise and consistent control. Despite the progress achieved by existing methods, ensuring fine-grained alignment with user intentions remains an open and challenging problem. In this work, we present Shape-for-Motion, a novel framework that incorporates a 3D proxy for precise and consistent video editing. Shape-for-Motion achieves this by converting the target object in the input video to a time-consistent mesh, i.e., a 3D proxy, allowing edits to be performed directly on the proxy and then inferred back to the video frames. To simplify the editing process, we design a novel Dual-Propagation Strategy that allows users to perform edits on the 3D mesh of a single frame, and the edits are then automatically propagated to the 3D meshes of the other frames. The 3D meshes for different frames are further projected onto the 2D space to produce the edited geometry and texture renderings, which serve as inputs to a decoupled video diffusion model for generating edited results. Our framework supports various precise and physically-consistent manipulations across the video frames, including pose editing, rotation, scaling, translation, texture modification, and object composition. Our approach marks a key step toward high-quality, controllable video editing workflows. Extensive experiments demonstrate the superiority and effectiveness of our approach. Project page: https://shapeformotion.github.io/

视频编辑3D代理扩散模型姿态控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。