无需标定即可实现红外与可见光图像的高精度超分辨率与融合
BeyondFusion: Self-Aligned Latent Diffusion for Calibration-Free Infrared Super-Resolution and Infrared-Visible Fusion

- 通过自对齐模块在潜空间中动态匹配多模态特征
- 在未校准条件下仍可重建高频红外细节并生成融合图像
- 适合移动设备上红外可见光协同感知场景
移动端红外-可见光成像通常将紧凑的红外传感器与高分辨率可见光相机结合以实现互补感知。然而,由于光学系统、视角、视场和曝光时间差异导致的跨传感器错位,严重制约了实际部署。本文提出BeyondFusion,一种统一的潜空间扩散框架,用于无标定条件下的可见光引导红外超分辨率与红外-可见光融合。该框架支持任务特定训练与联合训练,两任务共享同一生成过程的不同输出。不依赖显式配准或几何扭曲,BeyondFusion在去噪U-Net中引入跨模态自对齐(CMSA)模块,将红外与可见光潜变量重组至共享注意力空间,于去噪过程中学习内容自适应的跨模态对应关系。结合错位增强模块,模型可有效利用可见光结构与语义线索,同时保持热成像一致性,实现在未校准条件下的高频红外重建与信息丰富融合图像生成。在公开基准与移动端红外-可见光系统上的大量实验表明,其在对齐输入、低分辨率红外观测、合成错位及真实移动端非同步采集数据上均表现优异。消融实验、联合训练分析与下游行人检测任务进一步验证了其在无标定多模态成像中的有效性。
原文摘要 · Abstract (English)
Mobile infrared-visible imaging typically pairs a compact infrared sensor with a high-resolution visible camera for complementary perception. While cross-sensor misalignment caused by different optics, viewpoints, fields of view, and exposure timings hinders practical deployment. In this paper, we propose BeyondFusion, a unified latent diffusion framework for calibration-free visible-guided infrared super-resolution and infrared-visible fusion tasks. The proposed framework supports both task-specific training and joint training where two tasks are optimized and executed as two readouts of the same generative process. Instead of relying on explicit registration or geometric warping, BeyondFusion introduces a cross-modal self-aligning (CMSA) module into the denoising U-Net. CMSA reorganizes infrared and visible latent tokens into a shared attention space to learn content-adaptive cross-modal correspondence during the denoising process. Together with misalignment augmentation module, the model is facilitated to exploit visible structural and semantic cues while preserving thermal consistency, enabling high-frequency infrared reconstruction and informative fused-image generation under uncalibrated conditions. Extensive experiments on public benchmarks and a mobile infrared-visible imaging system show strong performance across aligned inputs, low-resolution infrared observations, synthetic misalignments, and real mobile captures with unsynchronized sensors. Ablation studies, unified training analysis, and downstream pedestrian detection further validate the effectiveness of BeyondFusion for calibration-free multimodal imaging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。