arXiv:2604.11140cs.CV2026-04

用稀疏超图和细粒度专家模型提升动态环境下的多模态目标检测精度

Hyper-FEOD: Sparse Hypergraph-Enhanced Frame-Event Object Detection with Fine-Grained MoE

  • 通过稀疏超图建模跨模态高阶关系,捕捉运动关键特征
  • 细粒度专家路由机制使不同区域特征得到针对性增强
  • 适合研究多模态感知、动态场景检测的开发者参考

将基于帧的RGB相机与事件流结合,是应对复杂动态条件下的鲁棒目标检测的有前景范式。然而,如何有效建模复杂的多模态交互,并解决RGB与事件数据间的语义异质性,仍是实现高精度检测的重大挑战。本文提出Hyper-FEOD框架,通过两个核心超图驱动组件协同强化跨模态表征学习。首先设计稀疏超图增强的跨模态融合(SHCF)模块,利用事件活动线索识别运动关键稀疏标记,并通过超图建模实现高阶关系推理,有效捕捉跨模态的复杂高阶依赖与丰富上下文关联。其次,开发面向异质语义需求的细粒度专家混合(FG-MoE)模块,通过具有不同超边连接模式的专用超图专家及空间门控机制,自适应地将特征路由至目标区域,实现精准增强。结合辅助路由器损失,该框架确保稳定端到端训练与最优特征优化。在广泛采用的RGB-Event基准测试中,Hyper-FEOD展现出卓越检测性能,显著优于现有最先进方法。

原文摘要 · Abstract (English)

The integration of frame-based RGB cameras with event streams constitutes a promising paradigm for robust object detection under challenging dynamic conditions. Nevertheless, effectively modeling intricate multi-modal interactions and reconciling the semantic heterogeneity between RGB and event data remain formidable challenges for high-precision detection. In this paper, we present Hyper-FEOD, a novel high-performance detection framework that synergistically strengthens cross-modal representation learning through two core hypergraph-driven components. Specifically, we first design a Sparse Hypergraph-enhanced Cross-Modal Fusion (SHCF) module that exploits event activity cues to identify motion-critical sparse tokens and performs high-order relational reasoning through hypergraph modeling. This design effectively captures intricate high-order dependencies and rich contextual correlations across modalities. Second, we develop a Fine-Grained Mixture-of-Experts (FG-MoE) module tailored to handle the heterogeneous semantic demands arising from distinct visual regions. By deploying specialized hypergraph experts with varying hyperedge connectivity pattern and incorporating a spatial gating mechanism, FG-MoE adaptively routes features to enable precise enhancement at target regions. Coupled with an auxiliary router loss, the proposed framework ensures stable end-to-end training and optimal feature refinement. Comprehensive experiments conducted on widely-adopted RGB-Event benchmarks show that Hyper-FEOD delivers superior detection performance and consistently outperforms existing state-of-the-art approaches by a notable margin.

多模态检测超图建模事件相机专家网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。