arXiv:2509.18693cs.CV2025-09

用视觉语言模型实现精准、可解释的遥感地表分类,不依赖具体类别名称。

MVT: Mask-Grounded Vision-Language Models for Taxonomy-Aligned Land-Cover Tagging

  • 用SAM2+领域自适应生成高精度区域掩码,不依赖类别标签。
  • 通过双阶段LoRA微调,让大模型生成符合分类体系的描述与标签。
  • 适合需要跨数据集泛化且重视可解释性的遥感分析任务。

遥感地表理解日益需要不依赖类别的系统,在跨数据集场景下保持空间精度和可解释性。我们研究在域偏移下的几何优先发现与解释设置,候选区域以类别无关方式划定,监督信号避免使用具体类别名称,改用匿名标识符。不同于开放集识别与开放世界学习,我们聚焦于将类别无关的掩码证据与分类体系对齐的场景解释,而非未知类拒绝或持续扩展。提出MVT三阶段框架:(i) 使用经过领域自适应的SAM2提取边界忠实的区域掩码;(ii) 通过多模态大模型的双阶段LoRA微调,实现掩码引导的语义标注与场景描述生成;(iii) 利用经分层专家评分校准的LLM-as-judge进行输出评估。在跨数据集分割迁移任务中(训练于OpenEarthMap,测试于LoveDA),领域自适应后的SAM2提升了掩码质量;双阶段多模态大模型微调显著提高了分类体系对齐标签的准确性与掩码引导描述的信息量。

原文摘要 · Abstract (English)

Land-cover understanding in remote sensing increasingly demands class-agnostic systems that generalize across datasets while remaining spatially precise and interpretable. We study a geometry-first discovery-and-interpretation setting under domain shift, where candidate regions are delineated class-agnostically and supervision avoids lexical class names via anonymized identifiers. Complementary to open-set recognition and open-world learning, we focus on coupling class-agnostic mask evidence with taxonomy-grounded scene interpretation, rather than unknown rejection or continual class expansion. We propose MVT, a three-stage framework that (i) extracts boundary-faithful region masks using SAM2 with domain adaptation, (ii) performs mask-grounded semantic tagging and scene description generation via dual-step LoRA fine-tuning of multimodal LLMs, and (iii) evaluates outputs with LLM-as-judge scoring calibrated by stratified expert ratings. On cross-dataset segmentation transfer (train on OpenEarthMap, evaluate on LoveDA), domain-adapted SAM2 improves mask quality; meanwhile, dual-step MLLM fine-tuning yields more accurate taxonomy-aligned tags and more informative mask-grounded scene descriptions.

遥感分类多模态掩码引导可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。