arXiv:2605.07574cs.CV2026-05

用偏振信息提升视觉语言模型对反光透明物体的识别能力

PolarVLM: Bridging the Semantic-Physical Gap in Vision-Language Models

论文配图:PolarVLM: Bridging the Semantic-Physical Gap in Vision-Language Models
图 1 · 摘自论文原文
  • 引入偏振物理参数,双流架构融合光学特性与视觉理解
  • 在反光识别上提升26.6%,玻璃计数提升34.0%
  • 构建首个偏振感知VQA基准,适合物理场景理解研究者

主流视觉语言模型(VLMs)因标准RGB输入的固有局限,难以应对反射、透明物体等光学模糊问题。偏振成像可捕捉极化物理参数以解决此类歧义,但现有方法受限于固定输出格式且缺乏开放推理能力。为此,我们提出PolarVLM,首个将偏振物理参数融入VLM的多模态框架。采用双流结构与渐进式两阶段训练策略,在避免物理误判的同时保持通用视觉能力。同时构建了首个偏振感知VQA基准PolarVQA,包含75,000个基于物理知识的指令微调样本,聚焦反射与透明场景。实验表明,PolarVLM在五个评估任务中整体超越RGB基线25.4%,在反射识别上提升26.6%,玻璃计数提升34.0%,成功实现物理感知语义理解。

原文摘要 · Abstract (English)

Mainstream vision-language models (VLMs) fundamentally struggle with severe optical ambiguities, such as reflections and transparent objects, due to the inherent limitations of standard RGB inputs. While polarization imaging captures polarimetric physical parameters that resolve these ambiguities, existing methods are constrained by fixed-format outputs and remain isolated from open-ended reasoning. To bridge this semantic-physical gap, we introduce PolarVLM, the first multimodal framework integrating polarimetric physical parameters into VLMs. By employing a dual-stream architecture and a progressive two-stage training strategy, PolarVLM effectively prevents physical misinterpretations while preserving general visual abilities. Complementing our architecture, we construct PolarVQA, the first benchmark for polarization-aware VQA, featuring 75K physics-grounded instruction-tuning pairs targeting reflective and transparent scenes. Experiments show that PolarVLM surpasses the RGB baseline by 25.4% overall across five evaluation tasks, with remarkable gains of 26.6% in reflection recognition and 34.0% in glass counting, successfully unlocking physics-aware semantic understanding.

视觉语言模型偏振成像物理感知多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。