arXiv:2605.06010cs.CVcs.AI2026-05被引 1

让视觉系统实时感知热信号,提升夜间和雾天的可靠性

Adding Thermal Awareness to Visual Systems in Real-Time via Distilled Diffusion Models

论文配图:Adding Thermal Awareness to Visual Systems in Real-Time via Distilled Diffusion Models
图 1 · 摘自论文原文
  • 用蒸馏扩散模型提取热图与可见光图的像素级差异作为监督信号
  • 在不同硬件上实现毫秒级推理速度,支持边缘设备部署
  • 无需重新训练即可接入现有视觉系统,适合自动驾驶等实时场景

纯RGB视觉模型在夜间和雾霾等复杂场景下表现不佳,导致性能下降与安全风险。红外成像可捕捉热源信息,提供关键补充,但现有高保真融合方法延迟过高,难以实现实时边缘部署。为此,我们提出FusionProxy,一种独立、即插即用的实时图像融合模块,具备扩散模型级别的质量。该方法利用教师样本集合的两种互补统计特性:原始图像空间中的像素级方差用于加权像素级监督,冻结主干网络内的像素级方差用于空间路由特征对齐。训练完成后,FusionProxy可直接集成到任意视觉感知系统中,无需联合优化。大量实验表明,该方法在静态识别任务中表现优异,在动态任务(如闭环自动驾驶)中显著提升鲁棒性。关键的是,FusionProxy在多种平台(从高端GPU到通用硬件)均实现实时推理,为全天候感知提供灵活且通用的解决方案。

原文摘要 · Abstract (English)

Purely RGB-based vision models often fail to provide reliable cues in challenging scenarios such as nighttime and fog, leading to degraded performance and safety risks. Infrared imaging captures heat-emitting sources and provides critical complementary information, but existing high-fidelity fusion methods suffer from prohibitive latency, rendering them impractical for real-time edge deployment. To address this, we propose FusionProxy, a real-time image fusion module designed as a fully independent, plug-and-play component with diffusion level quality. FusionProxy exploits two complementary statistics of a teacher sample ensemble: per-pixel variance in raw image space, used to weight pixel-level supervision, and per-pixel variance inside frozen foundation backbones, used to route feature-level alignment spatially. Once trained, FusionProxy can be directly integrated into any visual perception system without joint optimization. Extensive experiments demonstrate that our method achieves superior performance on static recognition tasks and significantly enhances robustness in dynamic tasks, including closed-loop autonomous driving. Crucially, FusionProxy achieves real-time inference speeds on diverse platforms, from high-end GPUs to commodity hardware, providing a flexible and generalizable solution for all-day perception.

视觉融合红外感知实时推理自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。