聚焦行人等难例目标,提升多车协同感知精度
FocalComm: Hard Instance-Aware Multi-Agent Perception
- 按难例动态提取关键特征,只传有用信息
- 在真实数据集上行人检测性能显著提升
- 适合关注交通安全的自动驾驶研究者
多智能体协同感知(CP)是提升自动驾驶安全性的有前景范式,尤其对行人等弱势道路使用者的3D感知具有重要意义。然而,现有方法通常以车辆检测指标为目标,对行人等小目标、高危对象表现不佳,检测失败可能引发严重后果。此外,以往方法依赖全特征交换,未能仅传递有助于降低漏检的关键特征。为此,本文提出FocalComm框架,专注于在互联智能体间交换面向难例的目标特征。FocalComm包含两项创新:(1) 可学习的渐进式难例挖掘(HIM)模块,用于每台设备提取难例相关特征;(2) 基于查询的特征级(中间层)融合机制,在协作过程中动态加权这些识别出的特征。实验表明,FocalComm在两个挑战性真实世界数据集(V2X-Real 和 DAIR-V2X)上,无论车辆中心还是基础设施中心设置,均优于当前最优协同感知方法。尤其在V2X-Real数据集上,行人检测性能取得显著提升。
原文摘要 · Abstract (English)
Multi-agent collaborative perception (CP) is a promising paradigm for improving autonomous driving safety, particularly for vulnerable road users like pedestrians, via robust 3D perception. However, existing CP approaches often optimize for vehicle detection performance metrics, underperforming on smaller, safety-critical objects such as pedestrians, where detection failures can be catastrophic. Furthermore, previous CP methods rely on full feature exchange rather than communicating only salient features that help reduce false negatives. To this end, we present FocalComm, a novel collaborative perception framework that focuses on exchanging hard-instance-oriented features among connected collaborative agents. FocalComm consists of two key novel designs: (1) a learnable progressive hard instance mining (HIM) module to extract hard instance-oriented features per agent, and (2) a query-based feature-level (intermediate) fusion technique that dynamically weights these identified features during collaboration. We show that FocalComm outperforms state-of-the-art collaborative perception methods on two challenging real-world datasets (V2X-Real and DAIR-V2X) across both vehicle-centric and infrastructure-centric collaborative setups. FocalComm also shows a strong performance gain in pedestrian detection in V2X-Real.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。