arXiv:2501.16803cs.ROcs.CV2025-01ICCV被引 2

提出轻量级跨模态注意力机制,提升多车协同感知精度与灵活性。

RG-Attn: Radian Glue Attention for Multi-modality Multi-agent Cooperative Perception

论文配图:RG-Attn: Radian Glue Attention for Multi-modality Multi-agent Cooperative Perception
图 1 · 摘自论文原文
  • 基于弧度约束注意力,通过坐标对齐实现跨模态特征融合
  • 在真实与仿真数据集上达到顶尖检测准确率,通信效率高
  • 支持多种传感器组合,适合实际自动驾驶部署

协同感知通过车路协同通信实现多智能体传感器融合,提升自动驾驶性能。然而,现有方法多依赖单模态数据共享,在包含激光雷达与摄像头的异构传感器设置下融合效果受限。为此,本文提出径向粘合注意力(RG-Attn),一种轻量且通用的跨模态融合模块,通过基于变换的坐标对齐和统一采样/逆策略,统一处理智能体内与智能体间融合。RG-Attn利用弧度注意力约束,列式操作几何一致区域,降低开销并保持空间一致性,实现精准鲁棒融合。基于此,设计三种协同架构:Paint-To-Puzzle(PTP)注重通信效率,要求所有智能体具备激光雷达;Co-Sketching-Co-Coloring(CoS-CoCo)提供最大灵活性,支持任意传感器配置(如仅激光雷达、仅相机或两者兼有),具备强跨模态泛化能力;Pyramid-RG-Attn Fusion(PRGAF)追求最高检测精度,计算开销较大。在模拟与真实数据集上的广泛评估表明,该框架在准确率、灵活性与效率方面均达到当前最优水平。

原文摘要 · Abstract (English)

Cooperative perception enhances autonomous driving by leveraging Vehicle-to-Everything (V2X) communication for multi-agent sensor fusion. However, most existing methods rely on single-modal data sharing, limiting fusion performance, particularly in heterogeneous sensor settings involving both LiDAR and cameras across vehicles and roadside units (RSUs). To address this, we propose Radian Glue Attention (RG-Attn), a lightweight and generalizable cross-modal fusion module that unifies intra-agent and inter-agent fusion via transformation-based coordinate alignment and a unified sampling/inversion strategy. RG-Attn efficiently aligns features through a radian-based attention constraint, operating column-wise on geometrically consistent regions to reduce overhead and preserve spatial coherence, thereby enabling accurate and robust fusion. Building upon RG-Attn, we propose three cooperative architectures. The first, Paint-To-Puzzle (PTP), prioritizes communication efficiency but assumes all agents have LiDAR, optionally paired with cameras. The second, Co-Sketching-Co-Coloring (CoS-CoCo), offers maximal flexibility, supporting any sensor setup (e.g., LiDAR-only, camera-only, or both) and enabling strong cross-modal generalization for real-world deployment. The third, Pyramid-RG-Attn Fusion (PRGAF), aims for peak detection accuracy with the highest computational overhead. Extensive evaluations on simulated and real-world datasets show our framework delivers state-of-the-art detection accuracy with high flexibility and efficiency. GitHub Link: https://github.com/LantaoLi/RG-Attn

协同感知跨模态融合自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。