arXiv:2512.06838cs.CV2025-12AAAI被引 3

用运动状态查询实现高效协同感知,解决视角不一下的通信与对齐难题。

SparseCoop: Cooperative Perception with Kinematic-Grounded Queries

  • 用带3D位置和速度的显式状态查询,精准对齐异步视角数据
  • 通过粗到精聚合提升融合鲁棒性,检测精度领先现有方法
  • 引入协同去噪训练任务,加速收敛且抗通信延迟

协同感知对自动驾驶至关重要,可克服单个车辆因遮挡和视域受限带来的问题。现有方法依赖密集的鸟瞰图(BEV)特征共享,面临通信成本随规模二次增长、异步或不同视角下对齐困难等问题。虽有稀疏查询方法出现,但普遍存在几何表示不足、融合策略不佳及训练不稳定等缺陷。本文提出SparseCoop,一种完全摒弃中间BEV表示的稀疏协同感知框架,用于3D目标检测与跟踪。核心创新包括:基于运动状态的实例查询(含3D位置与速度),实现精确时空对齐;粗到精聚合模块增强融合鲁棒性;协同实例去噪任务加速并稳定训练。在V2X-Seq和Griffin数据集上验证,SparseCoop达到当前最优性能,兼具高计算效率、低传输开销和强通信延迟鲁棒性。

原文摘要 · Abstract (English)

Cooperative perception is critical for autonomous driving, overcoming the inherent limitations of a single vehicle, such as occlusions and constrained fields-of-view. However, current approaches sharing dense Bird's-Eye-View (BEV) features are constrained by quadratically-scaling communication costs and the lack of flexibility and interpretability for precise alignment across asynchronous or disparate viewpoints. While emerging sparse query-based methods offer an alternative, they often suffer from inadequate geometric representations, suboptimal fusion strategies, and training instability. In this paper, we propose SparseCoop, a fully sparse cooperative perception framework for 3D detection and tracking that completely discards intermediate BEV representations. Our framework features a trio of innovations: a kinematic-grounded instance query that uses an explicit state vector with 3D geometry and velocity for precise spatio-temporal alignment; a coarse-to-fine aggregation module for robust fusion; and a cooperative instance denoising task to accelerate and stabilize training. Experiments on V2X-Seq and Griffin datasets show SparseCoop achieves state-of-the-art performance. Notably, it delivers this with superior computational efficiency, low transmission cost, and strong robustness to communication latency. Code is available at https://github.com/wang-jh18-SVM/SparseCoop.

协同感知稀疏查询3D检测自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。