arXiv:2506.19579cs.ROcs.AI2025-06

机器人用视觉语言模型看物品,发现3D打印物容易识别错误。

Fake or Real, Can Robots Tell? Evaluating VLM Robustness to Domain Shift in Single-View Robotic Scene Understanding

  • 用真实工具和3D打印复制品对比,测试模型在物理场景中的表现
  • 模型对3D打印物品的描述准确率显著下降,但对真实物品仍有效
  • 现有评测指标会误判错误描述,不适合评估机器人实际应用

机器人场景理解越来越多依赖视觉语言模型(VLM)生成环境的自然语言描述。本文系统评估机械臂单视角拍摄桌面上物体的命名任务,引入受控的物理领域迁移:将真实工具与几何结构相似但纹理、颜色、材质不同的3D打印复制品进行对比。我们在多个指标下测试了多款可本地部署的前沿VLM,评估语义对齐与事实准确性。结果表明,尽管模型能有效描述常见真实物体,但在3D打印物品上性能显著下降,即使形状熟悉。我们进一步揭示标准评估指标的关键缺陷:部分指标完全无法检测领域偏移,或奖励流畅但事实错误的描述。这些发现凸显了将基础模型应用于具身智能体的局限性,亟需更鲁棒的架构与评估协议用于实际机器人应用。

原文摘要 · Abstract (English)

Robotic scene understanding increasingly relies on Vision-Language Models (VLMs) to generate natural language descriptions of the environment. In this work, we systematically evaluate single-view object captioning for tabletop scenes captured by a robotic manipulator, introducing a controlled physical domain shift that contrasts real-world tools with geometrically similar 3D-printed counterparts that differ in texture, colour, and material. We benchmark a suite of state-of-the-art, locally deployable VLMs across multiple metrics to assess semantic alignment and factual grounding. Our results demonstrate that while VLMs describe common real-world objects effectively, performance degrades markedly on 3D-printed items despite their structurally familiar forms. We further expose critical vulnerabilities in standard evaluation metrics, showing that some fail to detect domain shifts entirely or reward fluent but factually incorrect captions. These findings highlight the limitations of deploying foundation models for embodied agents and the need for more robust architectures and evaluation protocols in physical robotic applications.

视觉语言模型机器人感知领域迁移评测漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。