首个统一模型,用语言提示搞定遥感图像去云、去 haze、融合等多任务修复。
A Unified Foundation Model for All-in-One Multi-Modal Remote Sensing Image Restoration and Fusion with Language Prompting
- 用最优传输对齐异源波段,三路专家网络分别处理空间、光谱与全局特征。
- 在百万级数据集上训练,11项任务均超越现有模型,跨任务迁移能力强。
- 适合遥感研究者快速适配新场景,尤其擅长处理多源异构影像数据。
遥感图像常受云雾、雾霾、噪声、分辨率限制及传感器差异影响。现有修复与融合方法针对每种退化类型需独立建模。本文提出语言条件下的大规模遥感图像修复模型 LLaRS,首个面向多模态、多任务的遥感低层视觉统一基础模型。LLaRS 采用 Sinkhorn-Knopp 最优传输将异源波段对齐至语义匹配槽,通过三类互补的专家混合层(卷积专家处理空间模式,通道混合专家保持光谱保真度,带低秩适配器的注意力专家捕捉全局上下文),并借助逐步动态权重调整稳定联合训练。为训练 LLaRS,构建了包含一百万样本的 LLaRS1M 数据集,覆盖十一项修复与增强任务,融合真实配对观测与可控合成退化,并加入多样自然语言提示。实验表明,LLaRS 持续优于七种竞争模型;参数高效微调实验验证其在未见数据上的强迁移能力与高适应效率。
原文摘要 · Abstract (English)
Remote sensing imagery suffers from clouds, haze, noise, resolution limits, and sensor heterogeneity. Existing restoration and fusion approaches train separate models per degradation type. In this work, we present Language-conditioned Large-scale Remote Sensing restoration model (LLaRS), the first unified foundation model for multi-modal and multi-task remote sensing low-level vision. LLaRS employs Sinkhorn-Knopp optimal transport to align heterogeneous bands into semantically matched slots, routes features through three complementary mixture-of-experts layers (convolutional experts for spatial patterns, channel-mixing experts for spectral fidelity, and attention experts with low-rank adapters for global context), and stabilizes joint training via step-level dynamic weight adjustment. To train LLaRS, we construct LLaRS1M, a million-scale multi-task dataset spanning eleven restoration and enhancement tasks, integrating real paired observations and controlled synthetic degradations with diverse natural language prompts. Experiments show LLaRS consistently outperforms seven competitive models, and parameter-efficient finetuning experiments demonstrate strong transfer capability and adaptation efficiency on unseen data. Repo: https://github.com/yc-cui/LLaRS
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。