测试视觉语言模型对非人形机器人的功能推理能力,发现其预测偏保守。
Assessing VLM-Driven Semantic-Affordance Inference for Non-Humanoid Robot Morphologies

- 构建真实与合成数据混合的机器人功能关系数据集
- 在多类物体和机器人形态下表现不一,误报率低但漏报率高
- 适合研究机器人功能推理与安全决策的学者参考
视觉语言模型(VLM)在理解人-物交互方面表现出色,但在非人形机器人上的应用仍鲜有研究。本文探讨了VLM是否能有效推断具有根本性不同身体形态的机器人对物体的功能感知,填补了该领域部署中的关键空白。我们提出一个新型混合数据集,融合真实世界机器人功能-物体关系标注与VLM生成的合成场景,并对多种物体类别和机器人形态下的VLM性能进行实证分析,揭示出功能推理能力存在显著差异。实验表明,尽管VLM在非人形机器人上展现出良好泛化能力,但其在不同物体领域的表现明显不一致。关键发现是:所有形态与物体类别中均呈现低误报率但高漏报率的稳定模式,说明VLM倾向于保守预测。这一现象在新型工具使用和非常规物体操作场景中尤为突出,提示在机器人系统中整合VLM需结合其他方法,以缓解过度保守行为,同时保留低误报带来的固有安全性优势。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have demonstrated remarkable capabilities in understanding human-object interactions, but their application to robotic systems with non-humanoid morphologies remains largely unexplored. This work investigates whether VLMs can effectively infer affordances for robots with fundamentally different embodiments than humans, addressing a critical gap in the deployment of these models for diverse robotic applications. We introduce a novel hybrid dataset that combines annotated real-world robotic affordance-object relations with VLM-generated synthetic scenarios, and perform an empirical analysis of VLM performance across multiple object categories and robot morphologies, revealing significant variations in affordance inference capabilities. Our experiments demonstrate that while VLMs show promising generalisation to non-humanoid robot forms, their performance is notably inconsistent across different object domains. Critically, we identify a consistent pattern of low false positive rates but high false negative rates across all morphologies and object categories, indicating that VLMs tend toward conservative affordance predictions. Our analysis reveals that this pattern is particularly pronounced for novel tool use scenarios and unconventional object manipulations, suggesting that effective integration of VLMs in robotic systems requires complementary approaches to mitigate over-conservative behaviour while preserving the inherent safety benefits of low false positive rates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。