arXiv:2507.13677cs.CVcs.AI2025-07中稿 · ITSC2025被引 5

解决车联网异构传感器融合难题,提升多模态协同感知精度

HeCoFuse: Cross-Modal Complementary V2X Cooperative Perception with Heterogeneous Sensors

  • 采用分层融合与注意力加权机制,自适应对齐跨模态特征
  • 在9种异构配置下保持21.74%~43.38%的3D mAP,最高达43.38%
  • 适用于车载与路侧设备混合部署场景,适合智能交通系统研究者

现实中的车路协同感知系统常因成本与部署差异面临异构传感器配置问题,导致特征融合困难与感知可靠性下降。为此,我们提出HeCoFuse,一个面向混合传感器配置(摄像头、激光雷达或两者兼有)的统一协同感知框架。通过引入分层融合机制,结合通道与空间注意力自适应加权特征,有效解决跨模态特征错位与表征质量不均问题。同时,设计自适应空间分辨率调节模块,在计算开销与融合效果间取得平衡。为进一步提升不同配置下的鲁棒性,引入动态调整融合方式的协同学习策略。在真实世界TUMTraf-V2X数据集上的实验表明,全传感器配置(LC+LC)下达到43.22% 3D mAP,优于CoopDet3D基线1.17%;在L+LC场景下更达43.38% 3D mAP,且在九种异构配置中保持21.74%至43.38%的3D mAP表现。该结果经CVPR 2025 DriveX挑战赛验证,为当前TUM-Traf V2X数据集最优表现。

原文摘要 · Abstract (English)

Real-world Vehicle-to-Everything (V2X) cooperative perception systems often operate under heterogeneous sensor configurations due to cost constraints and deployment variability across vehicles and infrastructure. This heterogeneity poses significant challenges for feature fusion and perception reliability. To address these issues, we propose HeCoFuse, a unified framework designed for cooperative perception across mixed sensor setups where nodes may carry Cameras (C), LiDARs (L), or both. By introducing a hierarchical fusion mechanism that adaptively weights features through a combination of channel-wise and spatial attention, HeCoFuse can tackle critical challenges such as cross-modality feature misalignment and imbalanced representation quality. In addition, an adaptive spatial resolution adjustment module is employed to balance computational cost and fusion effectiveness. To enhance robustness across different configurations, we further implement a cooperative learning strategy that dynamically adjusts fusion type based on available modalities. Experiments on the real-world TUMTraf-V2X dataset demonstrate that HeCoFuse achieves 43.22% 3D mAP under the full sensor configuration (LC+LC), outperforming the CoopDet3D baseline by 1.17%, and reaches an even higher 43.38% 3D mAP in the L+LC scenario, while maintaining 3D mAP in the range of 21.74% to 43.38% across nine heterogeneous sensor configurations. These results, validated by our first-place finish in the CVPR 2025 DriveX challenge, establish HeCoFuse as the current state-of-the-art on TUM-Traf V2X dataset while demonstrating robust performance across diverse sensor deployments.

车路协同多模态融合感知系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。