arXiv:2608.12515cs.CVcs.RO2026-08中稿 · ECCV

评测视觉语言模型从第一视角判断机器人危险距离,发现提示词可提升高危识别。

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images?

  • 用三种开源模型+提示策略+微调评估机器人第一视角图像的危险等级
  • 仅用提示词改进后,Qwen-VL在高危情况召回率显著提升
  • 模型标签正确但未必关注关键区域,说明空间理解仍不足

从机器人第一人称视角评估近身危险对人机环境中的安全具身导航至关重要,需结合视觉与上下文推理。我们评估了三个开源视觉语言模型(InternVL、Qwen-VL、SmolVLM)在将第一人称机器人图像分类为四个危险等级上的表现,对比三种提示策略和两轮QLoRA微调与分层随机基线。未微调时所有模型表现接近基线,微调仅带来适度整体提升。然而,采用高级提示的Qwen-VL在高危案例上的召回率显著高于其他模型。人物定位分析显示,危险分类正确并不对应更好空间定位,表明模型可能在未关注场景相关区域的情况下生成有效安全标签。结果表明当前视觉语言模型在细粒度近身行为推理与空间定位方面仍受限,尽管针对性提示与微调可在特定模型中改善高危检测。

原文摘要 · Abstract (English)

Assessing proxemic danger from a robot's egocentric perspective is critical for safe embodied navigation in human environments and requires both visual and contextual reasoning. We evaluate three opensource vision-language models (VLMs) (\textit{InternVL}, \textit{Qwen-VL}, and \textit{SmolVLM}) on the classification of egocentric robot images into four danger levels, comparing three prompting strategies and two rounds of QLoRA fine-tuning against a stratified random baseline. Without fine-tuning, all models perform near the baseline, while fine-tuning yields only modest overall improvements. However, \textit{Qwen-VL} with an advanced prompt achieves substantially higher recall for high-danger cases than the other models. An analysis of person localization further shows that correct danger classification does not correspond to better spatial grounding, indicating that a model may produce a useful safety label without attending to the relevant region of the scene. These results show that current VLMs remain limited in fine-grained proxemic reasoning and spatial grounding, although targeted prompting and fine-tuning can improve high-danger detection in selected models.

视觉语言模型机器人安全近身距离提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。