arXiv:2603.05754cs.RO2026-03中稿 · the 2026 IEEE/RSJ …被引 3

让机器人通过热成像感知未知环境,安全执行复杂操作。

Safe-Night VLA: Seeing the Unseen via Thermal-Perceptive Vision-Language-Action Models for Safety-Critical Manipulation

论文配图:Safe-Night VLA: Seeing the Unseen via Thermal-Perceptive Vision-Language-Action Models for Safety-Critical Manipulation
图 1 · 摘自论文原文
  • 融合热成像与视觉语言模型,实现基于热力学特性的语义理解。
  • 引入控制屏障函数,在分布外场景下确保操作安全,避免碰撞。
  • 适合需要在非结构化环境中安全作业的机器人应用,如医疗或救援。

当前视觉-语言-动作(VLA)模型主要依赖RGB感知,难以捕捉热信号等不可见模态。同时,端到端生成策略缺乏显式安全约束,在遇到障碍物或训练分布外的新场景时易失效。为此,我们提出Safe-Night VLA,一种多模态操作框架,使机器人能通过长波红外热感知看见“看不见”的信息,并在非结构化环境中实现热感知驱动的安全操作。该框架将长波红外热成像集成到预训练视觉语言主干中,实现基于热力学属性的语义推理;为确保分布外条件下安全执行,引入基于控制屏障函数的安全过滤器,实现在策略执行期间对工作空间约束的确定性保障。我们在Franka机械臂上进行了真实世界实验,提出了新的评估范式:温度条件操作、亚表面目标定位和反射混淆消解,且推理时保持约束执行。结果表明,Safe-Night VLA显著优于仅使用RGB的基线模型,证实基础模型可有效利用非可见物理模态实现鲁棒操作。

原文摘要 · Abstract (English)

Current Vision-Language-Action (VLA) models rely primarily on RGB perception, preventing them from capturing modalities such as thermal signals that are imperceptible to conventional visual sensors. Moreover, end-to-end generative policies lack explicit safety constraints, making them fragile when encountering obstacles and novel scenarios outside the training distribution. To address these limitations, we propose Safe-Night VLA, a multimodal manipulation framework that enables robots to see the unseen while enforcing rigorous safety constraints for thermal-aware manipulation in unstructured environments. Specifically, Safe-Night VLA integrates long-wave infrared thermal perception into a pre-trained vision-language backbone, enabling semantic reasoning grounded in thermodynamic properties. To ensure safe execution under out-of-distribution conditions, we incorporate a safety filter via control barrier functions, which provide deterministic workspace constraint enforcement during policy execution. We validate our framework through real-world experiments on a Franka manipulator, introducing a novel evaluation paradigm featuring temperature-conditioned manipulation, subsurface target localization, and reflection disambiguation, while maintaining constrained execution at inference time. Results demonstrate that Safe-Night VLA outperforms RGB-only baselines and provide empirical evidence that foundation models can effectively leverage non-visible physical modalities for robust manipulation.

机器人操作热成像安全控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。