arXiv:2606.25629cs.ROeess.SP2026-06中稿 · IROS 2026

让机器人在突发情况中快速决策,兼顾安全与实时性。

Event-Adaptive Motion Planning with Distilled Vision-Language Model in Safety-Critical Situations

论文配图:Event-Adaptive Motion Planning with Distilled Vision-Language Model in Safety-Critical Situations
图 1 · 摘自论文原文
  • 用轻量级视觉语言模型动态识别异常行为并触发响应
  • 通过策略级决策提升动态避障安全裕度,实测优于基线
  • 适合高风险场景下的机器人导航,如物流、医疗等

在安全关键型场景中,机器人导航面临未预见语义事件的挑战,碰撞主要源于动态障碍物的不可预测行为,而非未见物体。尽管大尺度视觉语言模型(VLM)具备出色的常识推理能力,但频繁调用其进行连续控制会引入严重计算延迟,破坏物理执行稳定性。为此,我们提出事件自适应运动规划(EAMP),一种高效的基于VLM的机器人导航框架。具体而言,提示可配置的语义事件触发器(PC-SET)持续监测短时序片段中的行为异常,一旦触发,便激活经物理验证语义蒸馏微调的事件触发式精简版SemNav-VLM,将检测到的异常映射为离散策略级决策。随后,语义模型预测控制(SMPC)模块将这些策略转化为优化目标与几何参考的动态重配置。在安全关键型物流场景中的大量实验表明,EAMP有效实现了高层推理与低层控制的对齐,在显著提升动态安全裕度的同时保持了实时效率。

原文摘要 · Abstract (English)

Robot navigation in safety-critical scenarios faces significant challenges from unforeseen semantic events, where collisions arise primarily from the unpredictable behaviors of dynamic agents rather than unseen objects. While large vision-language models (VLMs) offer remarkable capabilities in commonsense reasoning, frequently invoking them within the continuous control loop introduces severe computational latency, fundamentally destabilizing physical execution. To address these challenges, we propose event-adaptive motion planning (EAMP), an efficient framework for VLM-based robot navigation. Specifically, a prompt-configurable semantic event trigger (PC-SET) selectively activates semantic intervention by continuously monitoring short temporal clips for behavioral anomalies. Upon triggering, an event-triggered distilled SemNav-VLM, fine-tuned via physically verified semantic distillation, maps detected anomalies into discrete strategy-level decisions. Subsequently, a semantic model predictive control (SMPC) module translates these strategies into dynamic reconfigurations of optimization objectives and geometric references. Extensive experiments in safety-critical logistics scenarios demonstrate that EAMP effectively aligns high-level reasoning with low-level control, significantly improving dynamic safety margins over existing baselines while preserving real-time efficiency.

机器人导航视觉语言模型安全决策实时控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。