arXiv:2606.01362cs.GRcs.CV2026-06

用材质图统一实现视频中物体增删改,效果更自然。

AlbedoEdit: Unified Instance-Level Video Editing with Albedo Guidance

论文配图:AlbedoEdit: Unified Instance-Level Video Editing with Albedo Guidance
图 1 · 摘自论文原文
  • 用不变于光照的材质图控制细节编辑,用户友好
  • 支持物体插入、移除和纹理修改,三任务统一处理
  • 能自动模拟高光、阴影等真实视觉效果,适合创意应用

视频生成模型在合成逼真视频序列方面取得显著进展。但要拓展到更广泛的创造性下游应用,需实现细粒度的实例级视频编辑,包括物体插入、移除和纹理编辑,这已成为一个突出且具挑战性的问题。现有方法或仅提供粗略语义控制的统一生成框架,或为单一编辑任务设计专用框架,限制了跨场景的灵活性与适用性。为此,我们提出 AlbedoEdit,一种统一的生成式视频编辑框架,可同时支持物体插入、移除和纹理编辑。核心思想是利用不受光照影响的固有材质图(albedo map),该图不含高光、阴影和相互反射等干扰,可作为指定细粒度外观编辑的有效且直观机制。基于视频基础模型,AlbedoEdit 经微调后,将源 RGB 视频转化为编辑后的 RGB 视频,条件为用户编辑的第一帧材质图。在涵盖所有三种编辑任务的新配对合成数据集上训练,AlbedoEdit 隐式学习到对编辑内容的协调能力,并能模拟由编辑操作引发的复杂真实视觉效应,如镜面高光、柔影和镜面反射。AlbedoEdit 在定性和定量上均优于当前最先进方法。项目网页:https://vcai.mpi-inf.mpg.de/projects/AlbedoEdit/

原文摘要 · Abstract (English)

Video generative models have achieved remarkable progress in synthesizing photorealistic video sequences. However, enabling broader and more creative downstream applications requires fine-grained instance-level video editing, including object insertion, object removal, and texture editing, which has emerged as a prominent yet challenging problem. Existing approaches either propose unified generative frameworks with only coarse semantic control, or design task-specific frameworks for individual editing tasks, limiting their flexibility and applicability across diverse real-world scenarios. To address these limitations, we propose AlbedoEdit, a unified generative video editing framework that jointly supports object insertion, object removal, and texture editing. Our key insight is that the intrinsic albedo map, which is invariant to lighting and contains no specularity, shadowing and inter-reflection effects, provides an effective and user-friendly mechanism for specifying fine-grained appearance edits. Built upon video foundation models, AlbedoEdit is fine-tuned to translate source RGB videos into edited RGB videos, conditioned on a user-edited first-frame albedo. Trained on a new paired synthetic dataset covering all three editing tasks, AlbedoEdit implicitly learns to harmonize edited contents and simulate complex real-world visual effects triggered by editing operations, including specular highlights, soft shadows, and mirror reflections. AlbedoEdit demonstrates superior performance over state-of-the-art video editing approaches, both qualitatively and quantitatively. Project webpage is https://vcai.mpi-inf.mpg.de/projects/AlbedoEdit/.

视频编辑材质图生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。