通过智能拼图式传输,让多车感知通信量降500倍仍更准
JigsawComm: Joint Semantic Feature Encoding and Transmission for Communication-Efficient Cooperative Perception
- 用稀疏语义特征+效用评估,自动选最该传的信息
- 通信量减少20到500倍,准确率还比现有方法高
- 适合车联网、自动驾驶等带宽受限的协同场景
多智能体协同感知可突破单车感知的遮挡与范围限制,但受限于车辆与万物通信(V2X)带宽。现有方法通过压缩或启发式选消息提升效率,却忽略数据的语义相关性与跨智能体冗余。本文提出联合语义特征编码与传输问题,在通信预算下最大化感知精度,设计了端到端的语义感知框架JigsawComm,学习‘拼图’式多智能体特征传输。该框架使用正则化编码器提取稀疏、语义相关的特征,并引入轻量级特征效用估计器(FUE)预测各智能体每单元对下游感知任务的贡献。交换由FUE生成的紧凑元效用地图,用于在学习的效用代理下计算最优传输策略,天然消除跨智能体冗余,使特征传输量随智能体数增长保持在$O(1)$,元信息开销可忽略。整个流程通过可微调度模块端到端训练,使FUE与任务目标对齐。在OPV2V和DAIR-V2X基准上,JigsawComm将总数据量减少20至500倍,同时匹配或超越当前最优方法的准确率。
原文摘要 · Abstract (English)
Multi-agent cooperative perception (CP) promises to overcome the inherent occlusion and range limitations of single-agent systems in autonomous driving, yet its practicality is severely constrained by limited Vehicle-to-Everything (V2X) communication bandwidth. Existing approaches attempt to improve bandwidth efficiency via compression or heuristic message selection, but neglect the semantic relevance and cross-agent redundancy of the transmitted data. In this paper, we formulate a joint semantic feature encoding and transmission problem that maximizes CP accuracy under a communication budget, and introduce JigsawComm, an end-to-end semantic-aware framework that learns to ``assemble the puzzle'' of multi-agent feature transmission. JigsawComm uses a regularized encoder to extract \emph{sparse, semantically relevant features}, and a lightweight Feature Utility Estimator (FUE) to predict each agent's per-cell contribution to the downstream perception task. The FUE-generated compact meta utility maps are exchanged among agents and used to compute an optimal transmission policy under the learned utility proxy. This policy inherently \emph{eliminates cross-agent redundancy}, bounding the feature transmission payload to $\mathcal{O}(1)$ as the number of agents grows, while the meta information overhead remains negligible. The whole pipeline is trained end-to-end through a differentiable scheduling module, informing the FUE to be aligned with the task objective. On the OPV2V and DAIR-V2X benchmarks, JigsawComm reduces total data volume by over 20--500${\times}$ while matching or exceeding the accuracy of state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。