针对车联网中多类目标检测,提出自适应融合策略提升小大物体识别效果。
Class-Adaptive Cooperative Perception for Multi-Class LiDAR-based 3D Object Detection in V2X Systems

- 按物体类别动态调整特征提取与融合路径,匹配不同几何结构和点云密度。
- 在V2X-Real数据集上对卡车、行人检测性能显著提升,汽车也保持竞争力。
- 适合需要高精度多类目标感知的智能网联汽车系统部署场景。
协同感知使联网车辆与道路基础设施共享传感器数据,构建单个平台无法实现的融合场景表示。然而,现有协同3D目标检测器对所有物体类别采用统一融合策略,难以应对小物体与大物体在几何结构和点采样模式上的差异。这一问题还因评估协议狭窄而加剧,通常仅关注单一主导类别或少数协作设置,导致跨多样化车路协同场景的多类检测能力未被充分探索。为此,本文提出一种面向多类LiDAR 3D目标检测的类别自适应协同感知架构。模型集成四个组件:带可学习尺度路由的多尺度窗口注意力,用于空间自适应特征提取;类别特异性融合模块,将小物体与大物体分入不同的注意力融合路径;通过并行空洞卷积与通道重校准增强鸟瞰图上下文表示;以及类别平衡的目标加权机制,缓解高频类别的偏差。在V2X-Real基准上,实验覆盖以车辆为中心、基础设施为中心、车车、设施间及车路协同等五种设置,在相同主干网络与训练配置下验证。所提方法持续优于强基线中间融合模型,对卡车检测提升最大,对行人检测有明显改善,对汽车检测也达到竞争性水平。结果表明,将特征提取与融合与类别依赖的几何特性及点密度相匹配,能实现更均衡的现实车路协同感知。
原文摘要 · Abstract (English)
Cooperative perception allows connected vehicles and roadside infrastructure to share sensor observations, creating a fused scene representation beyond the capability of any single platform. However, most cooperative 3D object detectors use a uniform fusion strategy for all object classes, which limits their ability to handle the different geometric structures and point-sampling patterns of small and large objects. This problem is further reinforced by narrow evaluation protocols that often emphasize a single dominant class or only a few cooperation settings, leaving robust multi-class detection across diverse vehicle-to-everything interactions insufficiently explored. To address this gap, we propose a class-adaptive cooperative perception architecture for multi-class 3D object detection from LiDAR data. The model integrates four components: multi-scale window attention with learned scale routing for spatially adaptive feature extraction, a class-specific fusion module that separates small and large objects into attentive fusion pathways, bird's-eye-view enhancement through parallel dilated convolution and channel recalibration for richer contextual representation, and class-balanced objective weighting to reduce bias toward frequent categories. Experiments on the V2X-Real benchmark cover vehicle-centric, infrastructure-centric, vehicle-to-vehicle, infrastructure-to-infrastructure, and vehicle-to-infrastructure settings under identical backbone and training configurations. The proposed method consistently improves mean detection performance over strong intermediate-fusion baselines, with the largest gains on trucks, clear improvements on pedestrians, and competitive results on cars. These results show that aligning feature extraction and fusion with class-dependent geometry and point density leads to more balanced cooperative perception in realistic vehicle-to-everything deployments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。