用扩散模型将可见光图像转为热成像,解决机器人热成像数据少的问题。
ThermalDiffusion: Visual-to-Thermal Image-to-Image Translation for Autonomous Navigation
- 用条件扩散模型,通过自注意力学习物体热特性,生成真实感热图。
- 可为缺乏热成像的多模态自动驾驶数据集补充合成热图。
- 适合做自动驾驶感知、传感器融合的研究者和工程师参考。
自主系统依赖传感器感知环境,但相机、激光雷达和雷达在夜间或雾、霾、尘等恶劣条件下性能受限。热成像相机能通过物体的热信号提供有价值信息,尤其便于识别温度高于环境的人和车辆。本文聚焦热成像在机器人与自动化中的应用,核心挑战是热成像数据稀缺。现有多个用于自动驾驶研究的多模态数据集(如nuScenes、BDD100K)在场景分割、目标检测、深度估计等任务中广泛应用,但普遍缺少热成像数据。为此,本文提出利用条件扩散模型,基于自注意力机制学习真实世界物体的热特性,将现有RGB图像转换为合成热图像,从而高效扩充多模态数据集,推动热成像在自动驾驶中的快速部署。
原文摘要 · Abstract (English)
Autonomous systems rely on sensors to estimate the environment around them. However, cameras, LiDARs, and RADARs have their own limitations. In nighttime or degraded environments such as fog, mist, or dust, thermal cameras can provide valuable information regarding the presence of objects of interest due to their heat signature. They make it easy to identify humans and vehicles that are usually at higher temperatures compared to their surroundings. In this paper, we focus on the adaptation of thermal cameras for robotics and automation, where the biggest hurdle is the lack of data. Several multi-modal datasets are available for driving robotics research in tasks such as scene segmentation, object detection, and depth estimation, which are the cornerstone of autonomous systems. However, they are found to be lacking in thermal imagery. Our paper proposes a solution to augment these datasets with synthetic thermal data to enable widespread and rapid adaptation of thermal cameras. We explore the use of conditional diffusion models to convert existing RGB images to thermal images using self-attention to learn the thermal properties of real-world objects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。