arXiv:2509.06768cs.RO2025-09

让机器人用视觉语言模型实时识别危险并自动应对。

Embodied Hazard Mitigation using Vision-Language Models for Autonomous Mobile Robots

  • 融合视觉语言与大语言模型,实现多模态异常检测
  • 用户测试显示91.2%准确率,边缘计算延迟低
  • 适合需要主动避险的智能机器人应用

在动态环境中运行的自主机器人应能识别并报告异常。主动缓解机制可提升安全性和持续运行能力。本文提出一种多模态异常检测与缓解系统,整合视觉语言模型与大语言模型,实现实时识别城市与环境中的危险情境及冲突。该系统使机器人具备感知、理解、报告甚至在可能时响应异常的能力,通过主动检测机制和自动化缓解措施实现。关键贡献在于将危险状态与冲突状态融入机器人的决策框架,每类异常可触发特定缓解策略。用户研究(n=30)表明,系统在异常检测上达到91.2%的预测准确率,采用边缘人工智能架构实现了较低延迟的响应时间。

原文摘要 · Abstract (English)

Autonomous robots operating in dynamic environments should identify and report anomalies. Embodying proactive mitigation improves safety and operational continuity. This paper presents a multimodal anomaly detection and mitigation system that integrates vision-language models and large language models to identify and report hazardous situations and conflicts in real-time. The proposed system enables robots to perceive, interpret, report, and if possible respond to urban and environmental anomalies through proactive detection mechanisms and automated mitigation actions. A key contribution in this paper is the integration of Hazardous and Conflict states into the robot's decision-making framework, where each anomaly type can trigger specific mitigation strategies. User studies (n = 30) demonstrated the effectiveness of the system in anomaly detection with 91.2% prediction accuracy and relatively low latency response times using edge-ai architecture.

机器人视觉语言异常检测边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。