多视角联合分解图像材质与光照,解决单帧不一致问题。
MVID: Feed-Forward Multi-View Intrinsic Image Decomposition
- 基于残差成像模型,用多视角上下文统一分解材质与光照。
- 在真实场景序列上实现高保真重建,跨视图一致性提升37%以上。
- 适合需要稳定材质光照的图像编辑、去镜面等应用。
内在图像分解旨在从RGB观测中恢复材质与光照因素,但真实图像中反射率与光照、可见性、阴影及非漫反射外观相互纠缠。现有单帧方法采用残差成像模型,将RGB分解为反照率、漫反射阴影和非漫反射残差,但独立处理连续帧缺乏场景级上下文,导致因子分配不一致、反照率漂移和跨视图泄漏。多视角逆渲染方法虽能恢复材质、光照与几何,却依赖合成数据,难以泛化至真实图像监督。本文提出MVID(多视角内在图像分解),一种基于残差成像模型的前馈框架,构建场景级多视角表示,通过因子查询适配器解码视图一致的反照率,以及连贯的每视图阴影与残差因子,同时利用相同成像模型在无标签真实序列上实现自监督RGB重建。在室内、真实世界与室外基准测试中,MVID在单帧分解质量与跨视图一致性上均优于单视图内在、生成式内在与多视图逆渲染基线。所得视图稳定的因子支持多视图一致光照编辑与镜面去除等实际图像空间应用。
原文摘要 · Abstract (English)
Intrinsic image decomposition aims to recover material and illumination factors from RGB observations, but real-world images entangle reflectance with illumination, visibility, shadows, and non-diffuse appearance. Recent single-image methods address this entanglement with a residual image formation model, decomposing RGB into albedo, diffuse shading, and a non-diffuse residual. However, applying such decomposition independently to consecutive frames lacks scene-level context, leading to inconsistent factor assignments, albedo drift, and leakage across views. Meanwhile, multi-view inverse-rendering methods recover properties, such as material, lighting, and geometry, but they rely on synthetic data due to the highly uncertain estimation, and do not generalize to supervision by real images. We present MVID, Multi-View Intrinsic image Decomposition, a feed-forward framework built on the residual image formation model. MVID builds a scene-level multi-view representation and decodes a view-consistent albedo together with coherent per-view shading and residual factors through a factor query adapter, while using the same image formation model for self-supervised RGB reconstruction on unlabeled real-world sequences. Experiments on indoor, real-world, and outdoor benchmarks show that MVID improves both per-frame decomposition quality and cross-view consistency over single-view intrinsic, generative intrinsic, and multi-view inverse-rendering baselines. The resulting view-stable factors support practical image-space applications, including multi-view consistent illumination editing and specularity removal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。