用物理可微渲染探测视觉模型对3D场景的理解能力。
MRD: Using Physically Based Differentiable Rendering to Probe Vision Models for 3D Scene Understanding
- 通过可微物理渲染生成视觉上不同但模型响应相同的3D场景
- 模型对几何形状和材质的感知存在显著差异,但激活值高度相似
- 适合研究模型3D理解能力或人机视觉对比的学者使用
尽管深度学习在诸多视觉任务中表现优异,但其内部表征与决策机制仍难以解释。尽管视觉模型通常以2D输入训练,却常被认为隐式掌握了底层3D场景信息(如对部分遮挡具有鲁棒性,或能推理相对深度)。本文提出MRD(可微渲染的类异构体),利用基于物理的可微渲染,通过寻找在物理上不同但导致模型激活相同的3D场景参数,来探测模型对生成式3D场景属性的隐式理解。与以往基于像素的方法不同,该方法始终基于物理场景描述,可隔离物体形状、材质、光照等变量进行分析。作为初步验证,我们评估了多个模型在恢复几何形状和双向反射分布函数(BRDF)方面的表现。结果表明目标场景与优化后场景的模型激活高度相似,但视觉效果差异明显。定性分析揭示了模型对某些物理属性敏感或不变的特性。MRD为深入理解计算机与人类视觉提供了新工具,有助于分析物理场景参数如何驱动模型响应变化。
原文摘要 · Abstract (English)
While deep learning methods have achieved impressive success in many vision benchmarks, it remains difficult to understand and explain the representations and decisions of these models. Though vision models are typically trained on 2D inputs, they are often assumed to develop an implicit representation of the underlying 3D scene (for example, showing tolerance to partial occlusion, or the ability to reason about relative depth). Here, we introduce MRD (metamers rendered differentiably), an approach that uses physically based differentiable rendering to probe vision models' implicit understanding of generative 3D scene properties, by finding 3D scene parameters that are physically different but produce the same model activation (i.e. are model metamers). Unlike previous pixel-based methods for evaluating model representations, these reconstruction results are always grounded in physical scene descriptions. This means we can, for example, probe a model's sensitivity to object shape while holding material and lighting constant. As a proof-of-principle, we assess multiple models in their ability to recover scene parameters of geometry (shape) and bidirectional reflectance distribution function (material). The results show high similarity in model activation between target and optimized scenes, with varying visual results. Qualitatively, these reconstructions help investigate the physical scene attributes to which models are sensitive or invariant. MRD holds promise for advancing our understanding of both computer and human vision by enabling analysis of how physical scene parameters drive changes in model responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。