用VR+视觉语言模型帮救援机器人更清晰地告诉人类危险在哪
Autonomous VR-Based Risk Detection for Situational Awareness in Dangerous Settings

- 机器人在虚拟场景中用VLM识别危险并标注关键点
- 带标注的VR界面让用户感知更清晰,满意度更高
- 适合救援、安防等高危场景的人机协作研究
在灾难救援等高风险环境中,态势感知不仅依赖于危险检测,还取决于信息向操作人员的清晰传达。视觉语言模型(VLM)在安全关键场景的场景理解中展现出强大潜力,但其作为人机协作机器人系统组成部分的价值仍待探索。本文提出一种基于VR的人机交互框架,用于研究VLM辅助机器人如何支持模拟高危环境中的态势感知。在该系统中,机器人探索虚拟场景,通过VLM识别潜在危险并标注用户关注点,这些标注以沉浸式VR界面呈现给操作员。该框架实现了对机器人危险识别能力及安全信息传达效果的可控评估。研究结果表明,相比未标注的基线,带标注的VR界面更受青睐,参与者报告其具有高清晰度、实用性和舒适性。这表明将VLM驱动的机器人感知与沉浸式可视化结合,是提升高危环境下态势感知能力的有前景方法。
原文摘要 · Abstract (English)
In high-risk environments such as disaster response, situational awareness depends not only on detecting hazards but also on communicating them clearly to human operators. Vision Language Models (VLMs) have shown strong potential for scene understanding in safety-critical settings, yet their value as part of human-facing robotic systems remains underexplored. We present a VR-based Human Robot Interaction framework for studying how VLM-assisted robots can support situational awareness in simulated hazardous environments. In our system, a robot explores a virtual scene and queries a VLM to identify potential hazards and annotate user-facing points of interest. These annotations are presented to a human operator through an immersive VR interface. This framework enables controlled evaluation of both robotic hazard identification and the communication of safety-critical information to users. Results from our study indicate that the annotated VR interface was preferred over the unannotated baseline and that participants reported high clarity, usefulness, and comfort when interacting with the system. These findings suggest that combining VLM-based robotic perception with immersive visualization is a promising approach for supporting situational awareness in hazardous settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。