arXiv:2604.11671eess.SPcs.RO2026-04

用视觉语言模型和雷达物理特性融合识别材质,96%准确率无需训练。

VLMaterial: Vision-Language Model-Based Camera-Radar Fusion for Physics-Grounded Material Identification

论文配图:VLMaterial: Vision-Language Model-Based Camera-Radar Fusion for Physics-Grounded Material Identification
图 1 · 摘自论文原文
  • 双路架构:视觉提候选,雷达测介电常数
  • 雷达参数转为稳定语义参考,提升识别可信度
  • 自适应融合冲突,适合真实场景复杂材质识别

精准材质识别是智能感知系统安全高效交互物理世界的基础能力。例如,区分玻璃与塑料杯对安全至关重要,但视觉方法易受镜面反射、透明性及视觉欺骗影响。毫米波雷达在光照变化下仍具鲁棒性,但现有相机-雷达融合方法局限于封闭集类别且缺乏语义可解释性。本文提出VLMaterial,一种无需训练的框架,将视觉语言模型(VLM)与领域特定雷达知识融合,实现物理基础的材质识别。首先,设计双管道架构:光学管道使用Segment Anything模型与VLM生成材质候选;电磁表征管道通过有效峰值反射单元面积(PRCA)方法与加权向量合成,从雷达信号中提取内在介电常数。其次,采用上下文增强生成(CAG)策略,赋予VLM雷达特异性物理知识,使其能将电磁参数作为稳定参照。第三,引入自适应融合机制,基于不确定性估计解决跨模态冲突。我们在超过120次真实世界实验中评估了41种日常物品与4类视觉欺骗仿制品,覆盖多种环境。结果表明,VLMaterial达到96.08%的识别准确率,性能媲美最先进的封闭集基准,同时避免了大量任务特定数据收集与训练。

原文摘要 · Abstract (English)

Accurate material recognition is a fundamental capability for intelligent perception systems to interact safely and effectively with the physical world. For instance, distinguishing visually similar objects like glass and plastic cups is critical for safety but challenging for vision-based methods due to specular reflections, transparency, and visual deception. While millimeter-wave (mmWave) radar offers robust material sensing regardless of lighting, existing camera-radar fusion methods are limited to closed-set categories and lack semantic interpretability. In this paper, we introduce VLMaterial, a training-free framework that fuses vision-language models (VLMs) with domain-specific radar knowledge for physics-grounded material identification. First, we propose a dual-pipeline architecture: an optical pipeline uses the segment anything model and VLM for material candidate proposals, while an electromagnetic characterization pipeline extracts the intrinsic dielectric constant from radar signals via an effective peak reflection cell area (PRCA) method and weighted vector synthesis. Second, we employ a context-augmented generation (CAG) strategy to equip the VLM with radar-specific physical knowledge, enabling it to interpret electromagnetic parameters as stable references. Third, an adaptive fusion mechanism is introduced to intelligently integrate outputs from both sensors by resolving cross-modal conflicts based on uncertainty estimation. We evaluated VLMaterial in over 120 real-world experiments involving 41 diverse everyday objects and 4 typical visually deceptive counterfeits across varying environments. Experimental results demonstrate that VLMaterial achieves a recognition accuracy of 96.08%, delivering performance on par with state-of-the-art closed-set benchmarks while eliminating the need for extensive task-specific data collection and training.

材质识别多模态融合雷达感知VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。