用光影序列替代直接估计法,提升单目法线图的几何准确性
Monocular Normal Estimation via Shading Sequence Estimation
- 将法线估计转为光影序列预测,利用视频生成模型捕捉几何细节
- 在真实数据集上达到当前最佳性能,尤其对复杂物体更鲁棒
- 适合需要高精度表面重建的3D建模、逆向工程等场景
单目法线估计旨在从任意光照下的单张RGB图像中恢复法线图。现有方法依赖深度网络直接预测法线图,但常出现3D错位问题:尽管法线图外观合理,重建表面却无法对齐真实几何结构。我们指出,根源在于当前范式难以区分法线图中由细微颜色变化体现的差异性几何。为此,提出新范式——将法线估计重构为光影序列估计,因光影序列对几何变化更敏感。基于此,提出RoSE方法,利用图像到视频生成模型预测光影序列,并通过简单的最小二乘求解转换为法线图。为增强鲁棒性,RoSE在包含多样化形状、材质和光照条件的合成数据集MultiShade上训练。实验表明,RoSE在真实世界基准数据集上实现物体级单目法线估计的最先进性能。
原文摘要 · Abstract (English)
Monocular normal estimation aims to estimate the normal map from a single RGB image of an object under arbitrary lights. Existing methods rely on deep models to directly predict normal maps. However, they often suffer from 3D misalignment: while the estimated normal maps may appear to have a correct appearance, the reconstructed surfaces often fail to align with the geometric details. We argue that this misalignment stems from the current paradigm: the model struggles to distinguish and reconstruct varying geometry represented in normal maps, as the differences in underlying geometry are reflected only through relatively subtle color variations. To address this issue, we propose a new paradigm that reformulates normal estimation as shading sequence estimation, where shading sequences are more sensitive to various geometric information. Building on this paradigm, we present RoSE, a method that leverages image-to-video generative models to predict shading sequences. The predicted shading sequences are then converted into normal maps by solving a simple ordinary least-squares problem. To enhance robustness and better handle complex objects, RoSE is trained on a synthetic dataset, MultiShade, with diverse shapes, materials, and light conditions. Experiments demonstrate that RoSE achieves state-of-the-art performance on real-world benchmark datasets for object-based monocular normal estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。