首个专为遥感小目标设计的大模型,兼顾密集小物与大物检测
Bridging the Scale Gap: Balanced Tiny and General Object Detection in Remote Sensing Imagery
- 用动态路由融合多尺度特征,避免模型偏向大目标
- 根据密度自适应调整查询点数量和位置,提升资源利用效率
- 在多个数据集上实现小目标与通用目标的均衡高精度检测
遥感图像中的小目标检测近年来受到广泛关注。尽管取得进展,但在密集小目标与大目标共存场景下实现跨尺度平衡检测仍具挑战。尽管大模型已革新通用视觉任务,但其在遥感小目标检测中的应用尚未探索,主要受限于极端尺度差异和密度分布。为此,我们提出 ScaleBridge-Det,据知是首个专为小目标设计的大检测框架,通过尺度自适应专家路由与密度引导查询分配,在多样尺度间实现平衡性能。具体地,提出路由增强的混合注意力(REM)模块,通过自适应路由动态选择并融合特定尺度专家特征,缓解标准 MoE 模型对主导尺度的偏好,生成适用于小目标与大目标的互补且判别性多尺度表征。同时引入密度引导的动态查询(DGQ)模块,预测物体密度以自适应调节查询位置与数量,实现不同尺度对象的高效资源分配。该框架使 ScaleBridge-Det 在不牺牲性能的前提下,同时优化密集小目标与一般目标检测效果。在 AI-TOD-V2、DTOD 等基准数据集及跨域数据集 VisDrone 上的大量实验表明,ScaleBridge-Det 达到当前最优性能,并展现出优越的跨域鲁棒性。
原文摘要 · Abstract (English)
Tiny object detection in remote sensing imagery has attracted significant research interest in recent years. Despite recent progress, achieving balanced detection performance across diverse object scales remains a formidable challenge, particularly in scenarios where dense tiny objects and large objects coexist. Although large foundation models have revolutionized general vision tasks, their application to tiny object detection remains unexplored due to the extreme scale variation and density distribution inherent to remote sensing imagery. To bridge this scale gap, we propose ScaleBridge-Det, to the best of our knowledge, the first large detection framework designed for tiny objects, which could achieve balanced performance across diverse scales through scale-adaptive expert routing and density-guided query allocation. Specifically, we introduce a Routing-Enhanced Mixture Attention (REM) module that dynamically selects and fuses scale-specific expert features via adaptive routing to address the tendency of standard MoE models to favor dominant scales. REM generates complementary and discriminative multi-scale representations suitable for both tiny and large objects. Furthermore, we present a Density-Guided Dynamic Query (DGQ) module that predicts object density to adaptively adjust query positions and numbers, enabling efficient resource allocation for objects of varying scales. The proposed framework allows ScaleBridge-Det to simultaneously optimize performance for both dense tiny and general objects without trade-offs. Extensive experiments on benchmark and cross-domain datasets demonstrate that ScaleBridge-Det achieves state-of-the-art performance on AI-TOD-V2 and DTOD, while exhibiting superior cross-domain robustness on VisDrone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。