arXiv:2604.18866cs.CV2026-04

提出分层模块路由机制,提升航拍图像跨域目标检测能力

HMR-Net: Hierarchical Modular Routing for Cross-Domain Object Detection in Aerial Images

论文配图:HMR-Net: Hierarchical Modular Routing for Cross-Domain Object Detection in Aerial Images
图 1 · 摘自论文原文
  • 分两层路由:地理域与场景子区域分别匹配专家模块
  • 在4个数据集上实现跨域泛化与新类别检测性能提升
  • 无需微调即可通过文本描述识别未见物体类别

尽管目标检测技术不断进步,航拍图像仍具挑战性,因模型难以适应空间分辨率、场景构成和语义标签覆盖的差异。地理背景、传感器特性和物体分布的不同限制了传统模型学习一致可迁移表征的能力。通用方法在根本不同的域间强加统一表示,导致对特定区域内容表现差,且难以应对新类别。为此,我们提出一种新型模块化学习框架,实现航拍检测中的结构化专精。方法引入分层路由机制:第一层为域路由,利用潜在地理嵌入将输入分配给域专用专家模块;第二层为场景路由,将图像子区域分配至场景专用专家模块。该设计支持跨数据集及复杂场景内的专业化。此外,框架包含条件专家模块,可通过外部语义信息(如类别名或文本描述)在推理时识别新类别,无需重新训练或微调。相比单体表征,本方法为遥感目标检测提供了自适应框架。在四个数据集上的综合评估表明,该方法显著提升了多数据集泛化、区域级专精及开放类别检测性能。

原文摘要 · Abstract (English)

Despite advances in object detection, aerial imagery remains a challenging domain, as models often fail to generalize across variations in spatial resolution, scene composition, and semantic label coverage. Differences in geographic context, sensor characteristics, and object distributions across datasets limit the capacity of conventional models to learn consistent and transferable representations. Shared methods trained on such data tend to impose a unified representation across fundamentally different domains, resulting in poor performance on region-specific content and less flexibility when dealing with novel object categories. To address this, we propose a novel modular learning framework that enables structured specialization in aerial detection. Our method introduces a hierarchical routing mechanism with two levels of modularity: a domain routing layer that uses latent geographic embeddings to assign inputs to domain-specialized expert modules, and a scene routing mechanism that allocates image subregions to scene-specific expert modules. This allows our method to specialize across datasets and within complex scenes. Additionally, the framework contains a conditional expert module that uses external semantic information (e.g., category names or textual descriptions) to enable detection of novel object categories during inference, without the need for retraining or fine-tuning. By moving beyond monolithic representations, our method provides an adaptive framework for remote sensing object detection. Comprehensive evaluations on four datasets highlight improvements in multi-dataset generalization, region-level specialization, and open-category detection.

目标检测航拍图像模块化跨域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。