首个全量化多智能体协同感知系统,显著降低延迟与带宽开销。
QuantV2X: A Fully Quantized Multi-Agent System for Cooperative Perception
- 统一端到端量化策略,同时压缩模型计算与通信数据
- 低比特下仍达全精度精度,延迟降3.2倍,mAP30提升9.5
- 适合资源受限场景部署,支持更大模型在内存约束下运行
通过车与万物(V2X)通信实现协同感知,可有效缓解遮挡并扩展视野。然而,以往研究多关注精度提升,忽视效率、延迟与实际部署可行性。现有系统普遍依赖全精度模型,导致高计算与传输开销,难以在资源受限环境下实时运行。本文提出 extbf{QuantV2X},首个专为高效可扩展部署设计的多模态多智能体V2X协同感知全量化系统。该系统采用统一端到端量化策略,对神经网络模型与传输消息表示同步压缩,显著降低计算负载与通信带宽。值得注意的是,尽管在低比特约束下运行,QuantV2X仍达到与全精度系统相当的精度。更重要的是,在面向部署的评估指标下,其系统级延迟降低3.2倍,mAP30相比全精度基线提升9.5。此外,QuantV2X具备更强可扩展性,使更大更强大的模型能在严格内存预算内运行。这些结果证明了全量化多智能体中间融合系统的现实可行性。系统将公开发布以推动该领域研究:https://github.com/ucla-mobility/QuantV2X。
原文摘要 · Abstract (English)
Cooperative perception through Vehicle-to-Everything (V2X) communication offers significant potential for enhancing vehicle perception by mitigating occlusions and expanding the field of view. However, past research has predominantly focused on improving accuracy metrics without addressing the crucial system-level considerations of efficiency, latency, and real-world deployability. Noticeably, most existing systems rely on full-precision models, which incur high computational and transmission costs, making them impractical for real-time operation in resource-constrained environments. In this paper, we introduce \textbf{QuantV2X}, the first fully quantized multi-agent system designed specifically for efficient and scalable deployment of multi-modal, multi-agent V2X cooperative perception. QuantV2X introduces a unified end-to-end quantization strategy across both neural network models and transmitted message representations that simultaneously reduces computational load and transmission bandwidth. Remarkably, despite operating under low-bit constraints, QuantV2X achieves accuracy comparable to full-precision systems. More importantly, when evaluated under deployment-oriented metrics, QuantV2X reduces system-level latency by 3.2$\times$ and achieves a +9.5 improvement in mAP30 over full-precision baselines. Furthermore, QuantV2X scales more effectively, enabling larger and more capable models to fit within strict memory budgets. These results highlight the viability of a fully quantized multi-agent intermediate fusion system for real-world deployment. The system will be publicly released to promote research in this field: https://github.com/ucla-mobility/QuantV2X.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。