arXiv:2512.21003cs.CV2025-12被引 5

秒级完成多视角逆渲染,一次前向传播搞定材质光照一致性

MVInverse: Feed-forward Multi-view Inverse Rendering in Seconds

  • 通过跨视角交替注意力捕捉光照与材质一致性
  • 在真实场景视频上微调后,对野外图像泛化能力显著提升
  • 无需迭代优化,单次前向计算实现高精度几何材质重建

多视角逆渲染旨在从多视角图像中一致地恢复几何、材质和光照。现有单视角方法忽略跨视角关系,导致结果不一致;而多视角优化方法依赖缓慢的可微渲染与逐场景精调,计算成本高且难以扩展。为此,我们提出一种前馈式多视角逆渲染框架,直接从RGB图像序列预测空间变化的反照率、金属度、粗糙度、漫反射阴影和表面法线。通过跨视角交替注意力机制,模型同时捕获视图内长程光照交互与视图间材质一致性,实现单次前向传播下的场景级推理。由于真实数据稀缺,现有合成数据训练的模型在真实场景上泛化性差。为此,我们设计了一种基于一致性的微调策略,利用无标签真实视频增强多视角一致性与野外环境鲁棒性。在基准数据集上的大量实验表明,本方法在多视角一致性、材质与法线估计质量以及真实图像泛化性能上均达到当前最优水平。

原文摘要 · Abstract (English)

Multi-view inverse rendering aims to recover geometry, materials, and illumination consistently across multiple viewpoints. When applied to multi-view images, existing single-view approaches often ignore cross-view relationships, leading to inconsistent results. In contrast, multi-view optimization methods rely on slow differentiable rendering and per-scene refinement, making them computationally expensive and hard to scale. To address these limitations, we introduce a feed-forward multi-view inverse rendering framework that directly predicts spatially varying albedo, metallic, roughness, diffuse shading, and surface normals from sequences of RGB images. By alternating attention across views, our model captures both intra-view long-range lighting interactions and inter-view material consistency, enabling coherent scene-level reasoning within a single forward pass. Due to the scarcity of real-world training data, models trained on existing synthetic datasets often struggle to generalize to real-world scenes. To overcome this limitation, we propose a consistency-based finetuning strategy that leverages unlabeled real-world videos to enhance both multi-view coherence and robustness under in-the-wild conditions. Extensive experiments on benchmark datasets demonstrate that our method achieves state-of-the-art performance in terms of multi-view consistency, material and normal estimation quality, and generalization to real-world imagery. Project page: https://maddog241.github.io/mvinverse-page/

逆渲染多视角前馈模型真实泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。