用极化信息让小模型超越大模型,仅需4万张图就行
Revisiting Shape from Polarization in the Era of Vision Foundation Models
- 用1954个真实物体扫描建模,生成高质量极化数据集
- 仅4万样本就超越主流极化与纯RGB大模型
- 极化信号可大幅降低数据与参数需求,适合资源受限场景
我们发现,仅使用极化线索,一个在小数据集上训练的轻量级模型即可在单次拍摄的物体级表面法线估计任务中超越仅依赖RGB的视觉基础模型(VFMs)。虽然极化与表面几何存在强物理关联,长期以来被研究,但近期基于大规模数据和缩放定律的纯RGB VFMs性能大幅提升,已超越传统极化方法。这引发对极化是否仍有必要性的疑问,因其需专用硬件且训练数据有限。我们认为,以往极化方法表现较差并非源于极化模态本身,而是领域差距所致。主要来源有二:一是现有合成数据集使用有限且不真实的3D物体,几何简单、纹理随机,与真实形状不符;二是真实极化信号常受传感器噪声影响,而训练时未充分建模。为解决第一个问题,我们利用1,954个真实世界物体扫描结果渲染高质量极化数据集,并引入预训练DINOv3先验以提升对未知物体的泛化能力。为解决第二个问题,我们提出极化传感器感知的数据增强策略,更贴近真实条件。仅用4万训练场景,我们的方法显著超越当前最先进的极化方法及纯RGB VFMs。大量实验表明,极化线索可实现33倍数据量减少或8倍参数量减少,同时性能仍优于纯RGB模型。
原文摘要 · Abstract (English)
We show that, with polarization cues, a lightweight model trained on a small dataset can outperform RGB-only vision foundation models (VFMs) in single-shot object-level surface normal estimation. Shape from polarization (SfP) has long been studied due to the strong physical relationship between polarization and surface geometry. Meanwhile, driven by scaling laws, RGB-only VFMs trained on large datasets have recently achieved impressive performance and surpassed existing SfP methods. This situation raises questions about the necessity of polarization cues, which require specialized hardware and have limited training data. We argue that the weaker performance of prior SfP methods does not come from the polarization modality itself, but from domain gaps. These domain gaps mainly arise from two sources. First, existing synthetic datasets use limited and unrealistic 3D objects, with simple geometry and random texture maps that do not match the underlying shapes. Second, real-world polarization signals are often affected by sensor noise, which is not well modeled during training. To address the first issue, we render a high-quality polarization dataset using 1,954 3D-scanned real-world objects. We further incorporate pretrained DINOv3 priors to improve generalization to unseen objects. To address the second issue, we introduce polarization sensor-aware data augmentation that better reflects real-world conditions. With only 40K training scenes, our method significantly outperforms both state-of-the-art SfP approaches and RGB-only VFMs. Extensive experiments show that polarization cues enable a 33x reduction in training data or an 8x reduction in model parameters, while still achieving better performance than RGB-only counterparts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。