arXiv:2608.21099cs.CVcs.AI2026-08

让视觉与红外模态像社交伙伴一样协作,提升复杂环境下的目标检测能力

A2DINOv3: Rethinking Multi-Modal Object Detection via Socialized Collaboration

论文配图:A2DINOv3: Rethinking Multi-Modal Object Detection via Socialized Collaboration
图 1 · 摘自论文原文
  • 引入社会协作机制,让不同模态专家有选择地交流信息
  • 在四个数据集上均达顶尖性能,尤其在低光和红外场景优势明显
  • 适合做多模态感知、自动驾驶及安防监控系统的研究与开发

多模态目标检测对复杂环境(如弱光、恶劣天气)中的鲁棒场景理解至关重要。尽管近期视觉基础模型(如 DINOv3)表现出强大表征能力,但将其适配到多模态场景仍具挑战。现有密集跨模态融合策略常导致异构模态无差别交互,引入冗余信息并破坏预训练表征。为此,本文从社会学习视角出发,提出适配 DINOv3 的 A2DINOv3 框架,采用多专家协作结构与社会协作协议(SCP)。RGB 与红外分支作为异构专家,独立保留专长知识,通过受控、选择性交互交换互补信息,有效缓解有害的跨模态干扰,并防止预训练先验退化。此外,引入零初始化策略,逐步激活跨模态协作,实现从模态特异性学习到协同表征学习的平稳过渡。在四个多模态基准测试上——包括航拍检测(GAIIC)、自动驾驶(FLIR)、低光监控(LLVIP)及多样真实场景(M3FD)——A2DINOv3 均持续取得最先进性能。

原文摘要 · Abstract (English)

Multi-modal object detection is essential for robust scene understanding in challenging conditions, including low-light and adverse environments. Recent vision foundation models (e.g., DINOv3) have exhibited strong representation capabilities, yet adapting them to multi-modal scenarios remains challenging. Existing dense cross-modal fusion strategies often force heterogeneous modalities to interact indiscriminately, which may introduce redundant information and disrupt the valuable pre-trained representations. To address this issue, we revisit multi-modal fusion from the perspective of socialized learning and propose adapter to DINOv3 (A2DINOv3), a multi-expert collaboration framework with a Socialized Collaboration Protocol (SCP). Specifically, RGB and infrared branches are modeled as heterogeneous experts that independently preserve their specialized knowledge while exchanging complementary information through selective and constrained interactions. This design mitigates harmful cross-modal interference and prevents degradation of pre-trained priors during adaptation. Furthermore, a zero-initialization strategy is introduced to gradually activate cross-modal collaboration, enabling a smooth transition from modality-specific learning to cooperative representation learning. Extensive experiments on four multi-modal benchmarks, including aerial detection (GAIIC), autonomous driving (FLIR), low-light surveillance (LLVIP), and diverse real-world scenarios (M3FD), demonstrate that A2DINOv3 consistently achieves state-of-the-art performance in multi-modal object detection.

多模态检测视觉基础模型红外感知协作机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。