arXiv:2511.11270cs.CV2025-11

让视觉模型学会识别材料本质,提升对材质的感知能力。

Φeat: Physically Grounded Material Feature Representation

  • 用光照与几何变化下的同一材料对比预训练,强化材质特征敏感性。
  • 在材料分类和选择任务中显著优于传统模型,性能提升明显。
  • 适合需要精准材质理解的应用,如工业检测、虚拟现实。

尽管基础模型已成为通用视觉骨干网络,但其表征主要针对语义优化,缺乏对反射率等物理因素的显式建模,限制了其在需明确材质推理任务中的表现。我们提出Φeat,一种基于物理特性的新型材料感知视觉骨干网络,旨在增强对材料身份(包括反射率和微结构)的敏感性。不同于依赖通用数据增强的方法,我们通过对比同一材料在受控光照与几何变化下的观测进行预训练,使模型对非本质因素(如光照)具备不变性,同时保留对内在材质属性的敏感性。实验表明,该表征为以材料为中心的任务(如基于特征的材料选择与分类)提供了强大先验。结果证明,基于物理的弱监督是学习适配材质感知表征的有效策略。

原文摘要 · Abstract (English)

While foundation models have emerged as general-purpose visual backbones, their representations are primarily optimized for semantics and lack explicit modeling of physical factors, such as reflectance, hindering their efficacy in tasks requiring explicit material reasoning. We introduce $Φ$eat$, a novel material-grounded visual backbone that encourages a representation sensitive to material identity, including reflectance and mesostructure. Instead of relying on generic data augmentations, we pretrain our model by contrasting observations of the same material under controlled variations in lighting and geometry. This encourages invariance to extrinsic factors while preserving sensitivity to intrinsic material properties. We show that the resulting representation provides strong priors for material-centric tasks, including feature-based material selection and classification. Our results demonstrate that physically inspired weak supervision is an effective strategy for learning representations tailored to material perception.

材质表征视觉模型弱监督物理感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。