1秒内完成跨传感器影像融合,无需额外数据。
Training and Inference within 1 Second -- Tackle Cross-Sensor Degradation of Real-World Pansharpening with Efficient Residual Feature Tailoring
- 在特征层嵌入可高效训练的裁剪模块,解决跨传感器退化问题。
- 512×512图像处理仅需0.2秒,4000×4000图像3秒,比零样本方法快100倍以上。
- 适合实时遥感影像处理,尤其适用于无训练数据场景。
基于深度学习的影像融合方法进展迅速,但特定传感器预训练模型在其他传感器数据上泛化能力差。现有方法如重训练或零样本迁移存在耗时长或需额外数据的问题。本文提出一种模块化分解方法,识别出融合特征映射到最终图像通道空间的关键接口,并在此处引入物理感知的无监督损失训练的特征裁剪模块,实现跨传感器退化问题的特征级修复。方法采用分块训练与并行推理,显著提升效率。实验表明,该方法在多个真实数据集上达到最优性能与效率:512×512×8图像可在0.2秒内完成训练与推理,4000×4000×8图像仅需3秒(在RTX 3090 GPU上),相比零样本方法提速超100倍,且仅需部分测试输入,无需外部数据。
原文摘要 · Abstract (English)
Deep learning methods for pansharpening have advanced rapidly, yet models pretrained on data from a specific sensor often generalize poorly to data from other sensors. Existing methods to tackle such cross-sensor degradation include retraining model or zero-shot methods, but they are highly time-consuming or even need extra training data. To address these challenges, our method first performs modular decomposition on deep learning-based pansharpening models, revealing a general yet critical interface where high-dimensional fused features begin mapping to the channel space of the final image. % may need revisement A Feature Tailor is then integrated at this interface to address cross-sensor degradation at the feature level, and is trained efficiently with physics-aware unsupervised losses. Moreover, our method operates in a patch-wise manner, training on partial patches and performing parallel inference on all patches to boost efficiency. Our method offers two key advantages: (1) $\textit{Improved Generalization Ability}$: it significantly enhance performance in cross-sensor cases. (2) $\textit{Low Generalization Cost}$: it achieves sub-second training and inference, requiring only partial test inputs and no external data, whereas prior methods often take minutes or even hours. Experiments on the real-world data from multiple datasets demonstrate that our method achieves state-of-the-art quality and efficiency in tackling cross-sensor degradation. For example, training and inference of $512\times512\times8$ image within $\textit{0.2 seconds}$ and $4000\times4000\times8$ image within $\textit{3 seconds}$ at the fastest setting on a commonly used RTX 3090 GPU, which is over 100 times faster than zero-shot methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。