测试视觉语言模型对机器人碰撞的感知能力,发现现有模型可靠性不足。
Probing Collision Grounding in Vision-Language Models for Safe Human-Robot Collaboration

- 构建物理仿真基准,同步多视角图像与接触标签
- 最佳模型宏平均F1低于50%,难以准确判断碰撞风险
- 适合关注机器人安全监控与具身智能的研究者
安全人机协作不仅需要视觉描述,还需判断机器人是否安全分离、已发生碰撞或即将碰撞。我们提出碰撞接地(collision grounding)能力:将视觉观测与机器人几何、相机视角、场景布局、人体距离及运动趋势结合,以推断当前和潜在接触。为此,我们构建了基于Habitat 3.0的TouchSafeBench基准,包含2,940个室内共存模拟场景,涵盖社交导航与社交重排任务,提供同步多视角RGB-D数据、俯视轨迹图、校准相机元数据及模拟器生成的接触标签。研究两类实际部署任务:当前安全状态分类与碰撞前预警。在三个前沿或面向机器人应用的视觉语言模型及九种视觉表征上,当前模型仍不可靠:最佳平均宏F1低于50%;显式深度未自动转化为碰撞证据;机器人-场景接触识别持续难于人体接触风险评估。该基准揭示具身视觉语言模型的核心局限:视觉流利性不等于物理可问责性。可靠的安全监控系统需显式关联视角、机器人形态、度量几何与未来碰撞预测。
原文摘要 · Abstract (English)
Safe human--robot collaboration requires more than visual description: a monitor must determine whether the robot body is safely separated, already colliding with the scene or a person, or about to collide. We call this capability collision grounding: binding visual observations to robot body geometry, camera viewpoint, scene layout, human proximity, and temporal motion in order to infer present and imminent contact. We introduce TouchSafeBench, a physics-grounded benchmark for evaluating collision grounding in vision-language models (VLMs). Built in Habitat~3.0, TouchSafeBench contains 2,940 simulated indoor co-presence episodes across social navigation and social rearrangement, with synchronized multi-view RGB-D observations, top-down trajectory maps, calibrated camera metadata, and simulator-derived contact labels. We study two deployment-facing tasks: classifying the current safety state and warning about imminent collision before contact. Across three frontier or robotics-oriented VLMs and nine visual representations, current models remain far from reliable: the best average Macro-F1 stays below 50\%, explicit depth is not automatically transformed into robot-body collision evidence, and robot--scene contact is consistently harder than human-contact risk. TouchSafeBench reveals a central limitation of embodied VLMs: visual fluency does not imply physical accountability. Reliable robot safety monitors will need representations that explicitly bind viewpoint, robot morphology, metric geometry, and future collision. We will release the benchmark upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。