无需训练即可实现遥感图像开放词汇语义分割,效果超越现有方法。
Standalone DINOv3 for Training-Free Open-Vocabulary Semantic Segmentation in Remote Sensing

- 利用文本感知拉普拉斯传播与高斯点积上采样,提升分割精度。
- 在UDD5、DOTA、LoveDA上达到领先性能,大图处理能力更强。
- 适合无标注数据场景,为遥感智能分析提供新思路。
遥感语义分割受限于昂贵的像素级标注,推动了无需训练的开放词汇方法发展。DINOv3的发布带来了DINO.txt,使独立的DINO骨干网络具备图文对比学习能力,为开放词汇分割提供了可能。本文提出DinoSplat-OV,一种无需微调或额外预训练的训练自由框架,适配遥感图像密集分布、多尺度和大尺寸特性。设计两个核心模块:文本感知拉普拉斯传播模块通过结合文本语义相似性与局部视觉相似性,降低补丁级预测噪声,提升区域一致性并保留边界;高斯点积上采样模块通过RGB引导的各向异性聚合与测试时优化重建像素级特征。全局锚点滑动窗口策略支持大规模图像处理。在UDD5、DOTA、LoveDA上的实验表明,该方法性能优于或相当现有训练自由方法,有效填补了DINO系列模型在训练自由开放词汇分割中的空白,为该方向进一步发展提供可行路径。
原文摘要 · Abstract (English)
Remote sensing semantic segmentation is hindered by costly pixel-level annotations, motivating training-free open-vocabulary methods. Recently, the recent release of DINOv3 brings DINO.txt, which equips the standalone DINO backbone with image-text contrastive learning and thus opens up the possibility of open-vocabulary segmentation. We propose DinoSplat-OV, a training-free framework that adapts DINOv3 to remote sensing without fine-tuning or additional pretraining. Targeting the dense distribution, multi-scale nature, and large size of remote sensing imagery, we design two core modules. Its Text-aware Laplacian Propagation module de-noises patch-level predictions by combining textual semantic affinities with local visual similarity, improving regional consistency while preserving boundaries. Its Gaussian Splatting Upsampling module reconstructs pixel-level features through RGB-guided anisotropic aggregation and test-time optimization. A global-anchor sliding-window strategy further supports large-scale imagery. Experiments on UDD5, DOTA, and LoveDA demonstrate competitive or superior performance over existing training-free methods, effectively filling the gap of DINO-series models in training-free open-vocabulary segmentation and providing a viable new path for further advances in this direction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。