arXiv:2606.24225cs.CV2026-06

让视频中物体的移动旋转缩放更准确,保持阴影反射一致。

Geometry-Instructed Video Editing

论文配图:Geometry-Instructed Video Editing
图 1 · 摘自论文原文
  • 用深度框和朝向框统一描述物体3D状态变化
  • 生成前后对比视频对,实现几何编辑精准训练
  • 支持多种操作且在真实视频上表现良好

物体级几何编辑(如平移、旋转、缩放、复制或删除)是数字内容创作中的常规操作,但在生成式视频编辑中仍不可靠。核心挑战在于跨视角和时间准确指定目标物体的3D状态变化,并一致更新依赖几何的次级效应(如阴影和反射)。本文提出GIVE框架,通过统一的物体状态表示实现几何指令编辑。两个对齐视频的几何流分别描述编辑前后的物体:深度框编码粗略的3D位置与范围,朝向框提供与外观无关的方向提示。两者共同构成紧凑的前后几何规范。为学习这些编辑,构建可扩展的图形引擎流水线,执行物体级编辑程序并渲染受控的前后成对视频,隔离出意图的几何变换,同时保持次级效应一致性。实验表明,GIVE在统一框架下实现了忠实的几何编辑、时间连贯性及一致的次级效应,且在真实视频上展现出良好泛化能力。

原文摘要 · Abstract (English)

Object-level geometric edits, including translating, rotating, scaling, duplicating, or removing an object, are routine operations in digital content creation (DCC) workflows, yet they remain unreliable in generative video editing. The key challenge lies in specifying the target object's 3D state change unambiguously across viewpoint and time, while consistently updating geometry-dependent secondary effects such as shadows and reflections. We introduce GIVE, a geometry-instructed video editing framework that represents edits through a unified object-state formulation. Two video-aligned geometry streams describe the target object before and after editing: a depth-box encoding coarse 3D placement and extent, and an orientation-box providing an appearance-agnostic orientation cue. Together, these streams provide a compact pre/post geometric specification for object-state transitions. To provide paired supervision for learning these edits, we build a scalable graphics-engine pipeline that executes object-level edit programs and renders controlled before/after pairs, isolating the intended geometric edit while keeping secondary effects consistent with the transformation. Experimental results demonstrate that GIVE produces faithful geometric edits with temporal coherence and consistent secondary effects across operators in a unified framework, and shows promising transfer to in-the-wild videos. Project page: https://geometry-instructed-video-editing.github.io/give/

视频编辑几何建模生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。