arXiv:2503.23083cs.CVcs.AI2025-03被引 2

用高效微调技术让大模型在遥感视觉定位中表现更优且省算力

Efficient Adaptation For Remote Sensing Visual Grounding

  • 采用低秩适配(LoRA)等参数高效微调方法,仅更新少量参数
  • 在遥感视觉定位任务上达到或超过现有最优模型性能
  • 适合需要快速部署、资源受限的遥感多模态应用

预训练模型的适配已成为人工智能中一种高效策略,可显著降低从零训练模型的计算开销。在遥感领域,视觉定位(VG)研究仍较薄弱,该方法使强大视觉语言模型得以实现鲁棒的跨模态理解。本文将参数高效微调(PEFT)技术应用于遥感视觉定位任务:评估了在Grounding DINO中不同模块使用LoRA的效果,并采用BitFit与适配器对在通用视觉定位数据集上预训练的OFA模型进行微调。实验表明,该方法在性能上达到或超越当前最先进(SOTA)模型,同时大幅降低计算成本。本研究展示了PEFT技术在遥感多模态分析中的潜力,为无需全量训练提供了可行、经济的替代方案。

原文摘要 · Abstract (English)

Adapting pre-trained models has become an effective strategy in artificial intelligence, offering a scalable and efficient alternative to training models from scratch. In the context of remote sensing (RS), where visual grounding(VG) remains underexplored, this approach enables the deployment of powerful vision-language models to achieve robust cross-modal understanding while significantly reducing computational overhead. To address this, we applied Parameter Efficient Fine Tuning (PEFT) techniques to adapt these models for RS-specific VG tasks. Specifically, we evaluated LoRA placement across different modules in Grounding DINO and used BitFit and adapters to fine-tune the OFA foundation model pre-trained on general-purpose VG datasets. This approach achieved performance comparable to or surpassing current State Of The Art (SOTA) models while significantly reducing computational costs. This study highlights the potential of PEFT techniques to advance efficient and precise multi-modal analysis in RS, offering a practical and cost-effective alternative to full model training.

遥感视觉定位高效微调多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。