arXiv:2506.19303cs.RO2025-06中稿 · the 2025 Internati…被引 7

用视觉触觉融合模型,让机器人更准判断物体物理属性。

Robotic Perception with a Large Tactile-Vision-Language Model for Physical Property Inference

  • 将视觉与触觉信息融合进多模态语言模型,提升感知能力。
  • 在35种物体上测试,预测结果与真实测量高度相关。
  • 无需额外训练即可泛化到新物体,适合实际机器人应用。

准确推断物理属性能显著提升机器人操作能力,使其通过自适应抓取策略安全高效地处理物体。以往方法通常仅依赖触觉或视觉数据,难以全面捕捉物体特性。本文提出一种新型跨模态感知框架,将视觉观测与触觉表征融入多模态视觉-语言模型中。所提出的物理推理框架采用分层特征对齐机制和优化提示策略,使模型能够做出与真实测量高度相关的特定属性预测。在35种多样化物体上评估,该方法优于现有基线,并展现出强大的零样本泛化能力。

原文摘要 · Abstract (English)

Inferring physical properties can significantly enhance robotic manipulation by enabling robots to handle objects safely and efficiently through adaptive grasping strategies. Previous approaches have typically relied on either tactile or visual data, limiting their ability to fully capture properties. We introduce a novel cross-modal perception framework that integrates visual observations with tactile representations within a multimodal vision-language model. Our physical reasoning framework, which employs a hierarchical feature alignment mechanism and a refined prompting strategy, enables our model to make property-specific predictions that strongly correlate with ground-truth measurements. Evaluated on 35 diverse objects, our approach outperforms existing baselines and demonstrates strong zero-shot generalization. Keywords: tactile perception, visual-tactile fusion, physical property inference, multimodal integration, robot perception

机器人感知多模态融合触觉识别物理属性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。