arXiv:2606.05491cs.CVcs.RO2026-06中稿 · ICRA被引 1

无需配对图像即可融合可见光与热成像重建3D场景

Unpaired RGB-Thermal Gaussian-Splatting Using Visual Geometric Transformers

论文配图:Unpaired RGB-Thermal Gaussian-Splatting Using Visual Geometric Transformers
图 1 · 摘自论文原文
  • 用视觉几何变换器分别估计双模态相机位姿
  • 通过普鲁克斯特算法对齐位姿,实现无配对联合注册
  • 首次构建多模态一致性评估基准,适合跨模态重建研究者

结合可见光与热成像的多模态新视图合成(NVS)可实现高精度3D场景重建。然而现有方法通常依赖精确校准的可见光-热成像图像对或立体设置,限制了可扩展性与实际部署。为此,我们提出一种无配对可见光-热成像NVS框架,采用VGGT——一种3D前馈变压器架构,独立估计每种模态的相机位姿。通过普鲁克斯特算法与跨模态特征匹配器对齐位姿集,实现无需配对校准的联合注册。基于此对齐,进一步提出一种直接从无配对可见光与热图像中学习的多模态3D高斯点阵方法。在多样场景上的实验表明,该方法在热图像合成上表现优异,同时保持可见光保真度。此外,我们发现现有重建方法会产生缺乏跨模态一致性的模态特异性重建结果。因此,我们引入一个基准评估框架,用于严格评估各模态图像合成性能及重建场景的多模态一致性。

原文摘要 · Abstract (English)

Multi-modal novel view synthesis (NVS) combining RGB and thermal imagery enables precise 3D scene reconstruction with visual and thermal information. However, existing methods typically rely on precisely calibrated RGB-thermal image pairs or stereo setups, limiting scalability and practical deployment. To address this, we introduce a framework for unpaired RGB-thermal NVS that leverages VGGT, a 3D feed-forward transformer architecture, to independently estimate camera poses for each modality. The pose sets are then aligned using the Procrustes algorithm with a cross-modal feature matcher, enabling joint registration without paired calibration. Building on this alignment, we further propose a multi-modal 3D Gaussian Splatting approach that learns directly from unpaired RGB and thermal images. Experiments on diverse scenes demonstrate that our method achieves competitive performance in thermal view synthesis while maintaining RGB fidelity. Moreover, we show that existing reconstruction approaches can produce modality-specific reconstructions that lack cross-modal consistency. We thus introduce a benchmarking framework to rigorously evaluate both per-modality image synthesis and the multi-modal coherence of reconstructed scenes.

3D重建多模态高斯点阵热成像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。