用遮罩反馈减少车载感知数据传输量,提升云端大模型推理效率
CABLE: Cloud-Assisted Bandwidth-efficient LMM-based Encoding for V2X Systems

- 边缘端利用前一帧分割掩码与运动补偿生成感兴趣区域
- 仅上传感兴趣区域图像,通信量减少73%-87%,预填充速度提升5-8倍
- 适合车联网中带宽受限的实时感知系统部署
云端大型多模态模型可为车联网系统提供强大的开放词汇感知能力,但直接上传全分辨率画面会造成严重通信开销和高云端预填充延迟。本文提出CABLE框架,通过在边缘端利用自车运动补偿传播前一帧云分割掩码,结合残差运动线索进行细化,并通过通道包络合并断开区域,形成稳健的兴趣区域(ROI)。仅上传被ROI掩码覆盖的图像,同时将云端分割结果作为下一帧先验反馈,构建从掩码到ROI再到大模型的闭环。在nuScenes、WOD-ZB、Waymo、KITTI和CADC五个数据集上的实验表明,该方法在保持感知性能的前提下实现了显著通信节省,达到73%–87%的像素覆盖率降低,预计大模型预填充速度提升5–8倍,检测质量仅略有下降。
原文摘要 · Abstract (English)
Cloud-hosted large multimodal models (LMMs) can provide strong open-vocabulary perception for Vehicle-to-Everything systems, but naively transmitting full-resolution frames from edge to cloud causes severe communication overhead and high cloud-side prefill latency. We present CABLE, a cloud-assisted bandwidth-efficient LMM-based encoding framework for edge-cloud perception. CABLE propagates the previous cloud segmentation mask on the edge using ego-motion compensation, refines it with residual-motion cues, and consolidates disconnected regions via a corridor envelope to form a robust region of interest (ROI). Only ROI-masked images are uploaded, while the cloud segmentation output is fed back as the prior for the next frame, forming a mask-to-ROI-to-LMM feedback loop. Experiments on five datasets (nuScenes, WOD-ZB, Waymo, KITTI, and CADC) show consistent communication savings while largely preserving perception, achieving $73$--$87\%$ ROI pixel-coverage reduction with $5$--$8\times$ estimated LMM prefill speedup at a modest detection-quality trade-off relative to full-frame inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。