arXiv:2608.28517cs.CV2026-08

先学目标域先验,再做跨模态图像转换,提升真实感与细节保留。

Learning the Target Priors Before Image Translation: A Decoupled Training Paradigm for Cross-Modal Image Translation in Remote Sensing

论文配图:Learning the Target Priors Before Image Translation: A Decoupled Training Paradigm for Cross-Modal Image Translation in Remote Sensing
图 1 · 摘自论文原文
  • 先用无配对数据学目标域生成先验,再用少量配对数据做条件控制。
  • 在SAR转RGB和NIR转RGB任务中达到当前最优,仅需9.81%专用参数。
  • 适合数据稀缺的遥感图像转换,兼顾真实感与实例保真度。

遥感中的跨模态图像转换需在保留源域内容的同时匹配目标域分布。现有方法联合学习目标先验与跨模态依赖,忽视了关键不对称性:仅有后者依赖跨模态对应关系。本文通过条件得分与去噪风险分析形式化该差异,提出「先学目标先验再转换」(LTP-BIT)范式,解耦两项任务。首先利用大规模无配对影像学习目标域生成先验;随后冻结预训练主干,通过参数高效双流架构P-DART学习源条件控制。控制实验表明,先验匹配与规模扩展主要提升目标域真实感,而实例保真度更依赖条件适应。LTP-BIT在SAR-to-RGB与NIR-to-RGB基准上达最佳性能,仅使用9.81%任务特定参数。在QXS-SAROPT数据集上,仅用25%配对样本即保持近全数据集的实例保真度。

原文摘要 · Abstract (English)

Cross-modal image translation in remote sensing must preserve source-observed content while matching the target-domain distribution. Existing methods jointly learn the target prior and cross-modal dependence from scarce paired data, overlooking a key asymmetry: only the latter intrinsically requires cross-modal correspondence. We formalize this distinction through conditional-score and denoising-risk analyses and propose Learning the Target Priors Before Image Translation (LTP-BIT), a prior-first paradigm that decouples the two learning tasks. LTP-BIT first learns a target-domain generative prior from large-scale unpaired imagery, then retains the pretrained backbone weights and learns source-conditioned control through P-DART, a parameter-efficient dual-stream architecture. Controlled experiments show that prior matching and scaling primarily improve target-domain realism, whereas instance fidelity relies more strongly on conditional adaptation. LTP-BIT achieves state-of-the-art performance across SAR-to-RGB and NIR-to-RGB benchmarks using only 9.81% task-specific parameters. On QXS-SAROPT, it retains near-full-data instance fidelity with only 25% of the paired samples.

图像转换遥感先验学习参数效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。