arXiv:2509.21464cs.CVcs.RO2025-09中稿 · ICASSP 2026被引 1

用残差向量量化压缩多智能体感知数据,大幅降低通信开销。

Residual Vector Quantization For Communication-Efficient Multi-Agent Perception

  • 通过多阶段残差向量量化压缩特征,仅传输像素级码本索引。
  • 在6-30 bpp下实现1365倍压缩率,精度损失极小。
  • 适合自动驾驶等低带宽场景,利于V2X实际部署。

多智能体协同感知(CP)通过连接的智能体(如自动驾驶汽车、无人机、机器人)共享信息来提升场景理解能力,但通信带宽限制了其可扩展性。本文提出ReVQom,一种端到端的特征编码方法,通过简单瓶颈网络结合多阶段残差向量量化(RVQ),在保持空间身份的同时压缩中间特征。该方法仅需传输每像素的码本索引,将未压缩32位浮点特征的8192 bpp降至6-30 bpp。在真实世界DAIR-V2X CP数据集上,ReVQom在30 bpp时实现273倍压缩,在6 bpp时达1365倍。在18 bpp(455倍)下性能与原始特征感知相当或更优,6-12 bpp下支持超低带宽运行且性能平滑退化。ReVQom为高效准确的多智能体协同感知提供了实用路径,推动车联网(V2X)部署。

原文摘要 · Abstract (English)

Multi-agent collaborative perception (CP) improves scene understanding by sharing information across connected agents such as autonomous vehicles, unmanned aerial vehicles, and robots. Communication bandwidth, however, constrains scalability. We present ReVQom, a learned feature codec that preserves spatial identity while compressing intermediate features. ReVQom is an end-to-end method that compresses feature dimensions via a simple bottleneck network followed by multi-stage residual vector quantization (RVQ). This allows only per-pixel code indices to be transmitted, reducing payloads from 8192 bits per pixel (bpp) of uncompressed 32-bit float features to 6-30 bpp per agent with minimal accuracy loss. On DAIR-V2X real-world CP dataset, ReVQom achieves 273x compression at 30 bpp to 1365x compression at 6 bpp. At 18 bpp (455x), ReVQom matches or outperforms raw-feature CP, and at 6-12 bpp it enables ultra-low-bandwidth operation with graceful degradation. ReVQom allows efficient and accurate multi-agent collaborative perception with a step toward practical V2X deployment.

多智能体向量量化通信压缩自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。