arXiv:2601.09108cs.CV2026-01AAAI被引 3

用小参数动态引导大模型,提升遥感图像目标分割精度

Small but Mighty: Dynamic Wavelet Expert-Guided Fine-Tuning of Large-Scale Models for Optical Remote Sensing Object Segmentation

  • 设计动态小波专家模块,生成任务相关可训练特征
  • 在三个遥感数据集上超越21种前沿方法,显著提升分割性能
  • 适合需要高效微调大模型的遥感与医学图像分析场景

准确识别和分割光学遥感图像(ORSIs)中的目标对推动遥感应用至关重要。现有方法多基于中等规模预训练模型,采用多种优化策略实现全参数微调并取得良好效果。事实上,更深层、更大规模的基础模型能提供更强性能支持,但其海量参数导致直接全参数微调面临严重训练困难,如显存占用过高、计算成本激增,限制了大规模模型在现有工作中的探索。本文提出一种新型动态小波专家引导微调范式WEFT,仅需少量可训练参数,通过小波专家引导高效适配大模型至ORSIs分割任务。具体而言,引入任务特定的小波专家提取器,从多角度建模小波专家并动态调节输出,生成富含任务信息的可训练特征用于后续微调;进一步构建专家引导的条件适配器,先通过注入可训练特征增强冻结特征的细粒度感知,再迭代更新两类特征信息,实现高效微调。大量实验表明,WEFT不仅在三个ORSIs数据集上超越21种最先进方法,还在伪装、自然及医学场景中取得最优结果。

原文摘要 · Abstract (English)

Accurately localizing and segmenting relevant objects from optical remote sensing images (ORSIs) is critical for advancing remote sensing applications. Existing methods are typically built upon moderate-scale pre-trained models and employ diverse optimization strategies to achieve promising performance under full-parameter fine-tuning. In fact, deeper and larger-scale foundation models can provide stronger support for performance improvement. However, due to their massive number of parameters, directly adopting full-parameter fine-tuning leads to pronounced training difficulties, such as excessive GPU memory consumption and high computational costs, which result in extremely limited exploration of large-scale models in existing works. In this paper, we propose a novel dynamic wavelet expert-guided fine-tuning paradigm with fewer trainable parameters, dubbed WEFT, which efficiently adapts large-scale foundation models to ORSIs segmentation tasks by leveraging the guidance of wavelet experts. Specifically, we introduce a task-specific wavelet expert extractor to model wavelet experts from different perspectives and dynamically regulate their outputs, thereby generating trainable features enriched with task-specific information for subsequent fine-tuning. Furthermore, we construct an expert-guided conditional adapter that first enhances the fine-grained perception of frozen features for specific tasks by injecting trainable features, and then iteratively updates the information of both types of feature, allowing for efficient fine-tuning. Extensive experiments show that our WEFT not only outperforms 21 state-of-the-art (SOTA) methods on three ORSIs datasets, but also achieves optimal results in camouflage, natural, and medical scenarios. The source code is available at: https://github.com/CSYSI/WEFT.

遥感分割小波专家高效微调大模型适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。