arXiv:2504.19223cs.CVcs.AI2025-04被引 1

提出跨相机光谱图像表征学习模型,解决不同相机光谱差异带来的通用性难题。

CARL: Camera-Agnostic Representation Learning for Spectral Image Analysis

  • 设计自注意力与交叉注意力结合的光谱编码器,提取通用光谱特征。
  • 在医学、自动驾驶、卫星影像上验证,对真实和模拟跨相机差异均表现更优。
  • 适用于多模态光谱数据,适合作为未来光谱大模型的基础架构。

光谱成像在医疗、城市感知等领域具有广阔应用前景,是遥感中的关键模态。然而,不同光谱相机在通道数量和波长覆盖上的差异,导致现有AI方法依赖特定相机,泛化能力差。为此,我们提出CARL,一种面向RGB、多光谱和高光谱成像的相机无关表征学习模型。通过新型光谱编码器,结合自注意力与交叉注意力机制,将任意通道数的光谱图像转换为统一的相机无关表示。采用专为CARL设计的基于特征的自监督策略实现时空-光谱预训练。大规模实验在医疗影像、自动驾驶和卫星影像领域验证了该模型对光谱异质性的强鲁棒性,在模拟与真实跨相机光谱变化数据集上均表现领先。该方法具备可扩展性和多功能性,可作为未来光谱基础模型的核心架构。代码与模型权重已公开于https://github.com/IMSY-DKFZ/CARL。

原文摘要 · Abstract (English)

Spectral imaging offers promising applications across diverse domains, including medicine and urban scene understanding, and is already established as a critical modality in remote sensing. However, variability in channel dimensionality and captured wavelengths among spectral cameras impede the development of AI-driven methodologies, leading to camera-specific models with limited generalizability and inadequate cross-camera applicability. To address this bottleneck, we introduce CARL, a model for Camera-Agnostic Representation Learning across RGB, multispectral, and hyperspectral imaging modalities. To enable the conversion of a spectral image with any channel dimensionality to a camera-agnostic representation, we introduce a novel spectral encoder, featuring a self-attention-cross-attention mechanism, to distill salient spectral information into learned spectral representations. Spatio-spectral pre-training is achieved with a novel feature-based self-supervision strategy tailored to CARL. Large-scale experiments across the domains of medical imaging, autonomous driving, and satellite imaging demonstrate our model's unique robustness to spectral heterogeneity, outperforming on datasets with simulated and real-world cross-camera spectral variations. The scalability and versatility of the proposed approach position our model as a backbone for future spectral foundation models. Code and model weights are publicly available at https://github.com/IMSY-DKFZ/CARL.

光谱成像表征学习跨相机自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。