针对通信干扰下的车联网协同感知,提出自适应特征融合变压器模型。
AFFormer: Adaptive Feature Fusion Transformer for V2X Cooperative Perception under Channel Impairments

- 通过时序、多车和空间相关性建模,动态融合受损感知特征。
- 在理想与干扰条件下均优于现有方法,3D检测精度提升显著。
- 适合研究车联网感知鲁棒性或部署于复杂通信环境的工程师。
精准的3D目标检测对保障自动驾驶安全至关重要。车联网(V2X)协同感知通过共享感知数据提升检测能力,但易受噪声、衰落和干扰等信道失真影响。为增强智能交通系统的可靠性,本文提出自适应特征融合变压器(AFFormer),一种基于Transformer的框架,通过建模时间、跨车辆及空间相关性,缓解受损特征的负面影响。AFFormer引入三个关键模块:多车时序聚合(Multi-Agent and Temporal Aggregation)用于跨车与时间的上下文感知融合,双空间注意力(Dual Spatial Attention)高效建模空间依赖,不确定性引导融合(Uncertainty-Guided Fusion)实现熵驱动的特征精炼。此外,采用教师-学生知识蒸馏策略,通过早期协作监督对齐融合特征,进一步提升鲁棒性。在V2XSet与DAIR-V2X数据集上验证,AFFormer在理想与有损通信条件下均持续优于现有方法,展现出更强的抗通信损伤能力,并保持良好的效率-精度平衡。
原文摘要 · Abstract (English)
Accurate 3D object detection is essential for ensuring the safety of autonomous vehicles. Cooperative perception, which leverages vehicle-to-everything (V2X) communication to share perceptual data, enhances detection but is vulnerable to channel impairments, such as noise, fading, and interference. To strengthen the reliability of intelligent transportation systems, this work improves the robustness of V2X cooperative perception under communication conditions that reflect common channel impairments. This paper proposes an Adaptive Feature Fusion Transformer (AFFormer), a Transformer-based framework that mitigates the adverse effects of corrupted features by modeling temporal, inter-agent, and spatial correlations. AFFormer introduces three key modules: Multi-Agent and Temporal Aggregation for context-aware fusion across agents and over time, Dual Spatial Attention for efficient modeling of spatial dependencies, and Uncertainty-Guided Fusion for entropy-driven refinement of fused features. A teacher-student knowledge distillation strategy further enhances robustness by aligning fused features with reliable early-collaboration supervision. AFFormer is validated on the V2XSet and DAIR-V2X datasets, where it consistently outperforms existing methods under both ideal and impaired communication conditions, demonstrating improved robustness to communication-induced feature degradation while maintaining a competitive efficiency-accuracy trade-off.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。