arXiv:2603.12659cs.CV2026-03中稿 · CVPR

用教师模型指导学生模型,让视觉语言模型更好理解航拍图像。

AVION: Aerial Vision-Language Instruction from Offline Teacher to Prompt-Tuned Network

  • 用大模型生成文本原型,结合航拍图验证语义准确性。
  • 学生模型仅增加少量可学习提示,实现少样本分类准确率提升。
  • 适合需要轻量化适配遥感视觉语言模型的研究者使用。

将视觉语言模型适配至遥感影像仍面临两大挑战:文本表征语义覆盖有限,视觉特征适应性不足。这些问题在航拍场景中尤为突出,因图像呈现多样视觉特征且物体区分细微。我们提出 AVION,一个针对遥感适配的知識蒸馏框架。教师模块通过大语言模型收集描述,并利用遥感图像特征验证其有效性,构建语义丰富的文本原型。学生模块在视觉与语言编码器中引入轻量级可学习提示,由教师引导对齐嵌入及其跨模态关系。训练完成后,学生模型可在推理时独立运行。在六个光学遥感基准测试上,AVION 在少样本分类和基础类别准确率上均有提升,且不损害对新类别的泛化能力。同时显著提高跨模态检索的平均召回率,额外可训练参数极少。

原文摘要 · Abstract (English)

Adapting vision-language models to remote sensing imagery remains challenging due to two key factors: limited semantic coverage in textual representations and insufficient adaptability of visual features. These issues are particularly significant in aerial scenes, which involve various visual appearances and fine-grained object distinctions. We propose AVION, a knowledge distillation framework tailored for remote sensing adaptation of vision-language models. The teacher module constructs semantically rich textual prototypes by collecting descriptions from a large language model and verifying validity using remote sensing image features. The student module integrates lightweight and learnable prompts into both vision and language encoders, guided by the teacher to align embeddings and their cross-modal relationships. Once trained, the student operates independently during inference. Experiments on six optical remote sensing benchmarks show that AVION improves few-shot classification and base-class accuracy without degrading generalization to novel categories. It also enhances mean recall for cross-modal retrieval, with minimal additional trainable parameters.

遥感视觉语言知识蒸馏少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。