arXiv:2503.16426cs.CV2025-03被引 2

针对遥感图像目标稀疏问题,提出动态感知模型提升处理效率与精度

DynamicVis: Dynamic Visual Perception for Efficient Remote Sensing Foundation Models

  • 设计动态区域感知机制,仅处理高显著性区域,跳过冗余背景
  • 在百万级数据上预训练,显式分离稀疏目标与密集背景,提升定位能力
  • 适用于小目标检测、变化检测等稀疏目标感知任务,性能显著领先

遥感技术发展推动了高分辨率地球观测,但利用现代视觉基础模型(VFM)解析此类图像仍面临挑战。与以物体为中心的自然图像不同,遥感图像具有极端目标稀疏性和巨大空间冗余特征:关键目标(如船只、车辆)通常占据不到1%的空间,被广阔无目标背景包围。现有模型多依赖均匀密集计算(如ViT)和像素重建预训练范式(如MAE),在冗余背景上浪费大量算力,并弱化了稀疏目标的特征表示。为此,我们提出DynamicVis,一种专为遥感图像稀疏特性设计的视觉基础模型。其架构引入动态区域感知状态空间模型(SSM),自适应路由并逐步建模任务相关高显著性标记,同时采用无参数整合方式处理背景上下文,将超长二维标记序列(约10万)的计算复杂度大幅降低。关键在于,我们提出新型区域级元嵌入多实例学习(MIL)预训练范式,在百万规模数据集上训练,显式在潜在语义空间中分离稀疏前景与密集背景,克服传统像素重建方法的语义模糊性。在九个多样化下游任务上的广泛评估表明,DynamicVis表现出卓越效能,尤其在稀疏目标与实例级感知任务(如小目标检测、变化检测)中占据主导地位。

原文摘要 · Abstract (English)

The advancement of RS technology has enabled high-resolution Earth observation; however, interpreting these images using modern VFMs remains a significant challenge. Unlike object-centric natural images, RS imagery is fundamentally characterized by extreme target sparsity and massive spatial redundancy. Key objects of interest (e.g., ships, vehicles) often occupy less than 1% of the spatial extent, surrounded by vast, target-free backgrounds. Existing VFMs predominantly rely on uniform dense processing (e.g., ViTs) and pixel-reconstruction pre-training paradigms (e.g., MAE). These approaches inherently waste substantial computational capacity on modeling redundant backgrounds and inadvertently dilute the feature representations of small, sparse targets. To bridge this structural misalignment, we propose DynamicVis, a visual foundation model explicitly tailored to the sparse nature of RS imagery. Architecturally, DynamicVis introduces a Dynamic Region-Aware SSM that bypasses uniform computation. It adaptively routes and incrementally models only task-relevant, high-salience tokens while employing a parameter-free integration for background context, drastically reducing the complexity of processing ultra-long 2D token sequences ($\sim$100,000). Crucially, to equip the network with robust spatial-selection capabilities, we propose a novel Region-Level Meta-Embedding Multi-Instance Learning (MIL) pre-training paradigm. Trained on a million-scale dataset, this paradigm explicitly disentangles sparse foreground instances from dense backgrounds in the latent semantic space, overcoming the semantic ambiguity of conventional pixel-reconstruction methods. Extensive evaluations across nine diverse downstream tasks reveal that DynamicVis exhibits exceptional efficacy, particularly dominating in sparse-target and instance-level perception tasks (e.g., small object detection, and change detection).

遥感图像动态感知小目标检测高效模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。