解决多智能体感知中的时间延迟与噪声问题,提升复杂交通下的协同感知能力。
CATNet: Collaborative Alignment and Transformation Network for Cooperative Perception
- 通过时序递归同步机制对齐异步特征流,构建统一时空表示空间。
- 设计双分支小波去噪器,有效抑制全局噪声并修复局部特征失真。
- 动态选择关键感知特征,增强融合鲁棒性,适合自动驾驶等实际场景。
协同感知通过整合多个智能体的互补信息显著提升场景理解能力。然而,现有研究常忽视真实多源数据融合中的关键挑战,即高时间延迟和多源噪声。为此,本文提出协作对齐与转换网络(CATNet),一种自适应补偿框架,用于解决多智能体系统中的时间延迟与噪声干扰。核心创新包括:首先,引入时空递归同步(STSync)机制,通过相邻帧差异建模对齐异步特征流,建立时空统一表示空间;其次,设计双分支小波增强去噪器(WTDen),在对齐表示中抑制全局噪声并重构局部特征失真;第三,构建自适应特征选择器(AdpSel),动态聚焦关键感知特征以实现鲁棒融合。在多个数据集上的大量实验表明,CATNet在复杂交通条件下持续优于现有方法,验证了其出色的鲁棒性与适应性。
原文摘要 · Abstract (English)
Cooperative perception significantly enhances scene understanding by integrating complementary information from diverse agents. However, existing research often overlooks critical challenges inherent in real-world multi-source data integration, specifically high temporal latency and multi-source noise. To address these practical limitations, we propose Collaborative Alignment and Transformation Network (CATNet), an adaptive compensation framework that resolves temporal latency and noise interference in multi-agent systems. Our key innovations can be summarized in three aspects. First, we introduce a Spatio-Temporal Recurrent Synchronization (STSync) that aligns asynchronous feature streams via adjacent-frame differential modeling, establishing a temporal-spatially unified representation space. Second, we design a Dual-Branch Wavelet Enhanced Denoiser (WTDen) that suppresses global noise and reconstructs localized feature distortions within aligned representations. Third, we construct an Adaptive Feature Selector (AdpSel) that dynamically focuses on critical perceptual features for robust fusion. Extensive experiments on multiple datasets demonstrate that CATNet consistently outperforms existing methods under complex traffic conditions, proving its superior robustness and adaptability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。