arXiv:2603.25044cs.RO2026-03被引 1

让机器人通过热成像感知环境,提升安全与任务成功率

ThermoAct:Thermal-Aware Vision-Language-Action Models for Robotic Perception and Decision-Making

  • 用视觉-语言-动作框架融合热成像数据,实现智能决策
  • 实测表明新系统任务成功率和安全性优于纯视觉方案
  • 适合需高安全性的工业协作机器人场景

在人机协作环境中,越来越多研究关注整合视觉以外的多源传感器数据,以实现更安全、智能的任务执行。尽管热成像数据对提升机器人安全性和运行效率至关重要,但以往研究对其集成关注较少。本文提出一种新型视觉-语言-动作(VLA)框架,将热成像信息融入机器人任务执行中。该系统利用视觉-语言模型(VLM)作为高层规划器,解析复杂自然语言指令并分解为简单子任务,提升数据采集效率与复杂操作的推理能力。与仅依赖视觉数据的传统方法不同,本方法融合热成像信息,使机器人能够感知物理特性并主动保障环境安全。真实场景实验验证了该框架的可行性,结果表明其相比现有视觉系统具有更高的任务成功率和安全性。

原文摘要 · Abstract (English)

In recent human-robot collaboration environments, there is a growing focus on integrating diverse sensor data beyond visual information to enable safer and more intelligent task execution. Although thermal data can be crucial for enhancing robot safety and operational efficiency, its integration has been relatively overlooked in prior research. This paper proposes a novel Vision-Language-Action (VLA) framework that incorporates thermal information for robot task execution. The proposed system leverages a Vision-Language Model (VLM) as a high-level planner to interpret complex natural language commands and decompose them into simpler sub-tasks. This approach facilitates efficient data collection and robust reasoning for complex operations. Unlike conventional methods that rely solely on visual data, our approach integrates thermal information, enabling the robot to perceive physical properties and proactively ensure environmental safety. Experimental results from real-world task scenarios validate the feasibility of our proposed framework, suggesting its potential to enhance task success rates and safety compared to existing vision-based systems.

机器人感知多模态热成像智能决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。