arXiv:2503.01107cs.CV2025-03CVPR被引 13

用生成模型实现视频中3D物体位置的时序一致编辑

VideoHandles: Editing 3D Object Compositions in Videos Using Video Generative Priors

  • 通过共享3D重建结构,将生成模型特征升维后编辑再投影回帧
  • 可在静态场景视频中统一调整物体3D位置,保持时间一致性
  • 无需训练,适合需要快速修改物体布局的视觉创作人群

图像和视频编辑的生成方法利用生成模型作为先验,在信息不完整的情况下进行编辑,例如改变单张图像中3D物体的构图。近期方法在图像编辑上取得良好效果,但在视频领域,现有方法主要关注物体外观、运动或相机运动的编辑,尚无针对视频中物体构图的编辑方法。本文提出一种新方法,用于在包含相机运动的静态场景视频中编辑3D物体的构图。该方法可跨所有帧以时序一致的方式编辑3D物体的位置。其核心是将生成模型的中间特征提升至共享的3D重建空间,编辑该重建,再将特征投影回各帧。据我们所知,这是首个基于生成先验的视频物体构图编辑方法。本方法简单且无需训练,性能优于当前最先进的图像编辑基线。

原文摘要 · Abstract (English)

Generative methods for image and video editing use generative models as priors to perform edits despite incomplete information, such as changing the composition of 3D objects shown in a single image. Recent methods have shown promising composition editing results in the image setting, but in the video setting, editing methods have focused on editing object's appearance and motion, or camera motion, and as a result, methods to edit object composition in videos are still missing. We propose \name as a method for editing 3D object compositions in videos of static scenes with camera motion. Our approach allows editing the 3D position of a 3D object across all frames of a video in a temporally consistent manner. This is achieved by lifting intermediate features of a generative model to a 3D reconstruction that is shared between all frames, editing the reconstruction, and projecting the features on the edited reconstruction back to each frame. To the best of our knowledge, this is the first generative approach to edit object compositions in videos. Our approach is simple and training-free, while outperforming state-of-the-art image editing baselines.

视频编辑3D重建生成模型时序一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。