用光流驱动语义通信,大幅降低3D目标检测传输数据量
Task-Oriented Semantic Communication for Stereo-Vision 3D Object Detection
- 基于光流联合提取左右图像语义,减少对齐信息损失
- 在关键区域增强目标周围语义传输,提升深度估计精度
- 适用于低信噪比场景,适合边缘计算与远程检测任务
随着计算机视觉的发展,3D目标检测在诸多实际应用中愈发重要。受限于传感器端硬件算力,检测任务常部署于远程计算设备或云端执行复杂算法,带来巨大的数据传输开销。为此,本文提出一种面向立体视觉3D目标检测的光流驱动语义通信框架。该框架充分利用立体视觉3D检测对图像语义信息的高度依赖性,优先传输关键语义信息,从而在保证检测精度的前提下显著减少总传输数据量。具体而言,设计光流驱动模块,联合从左右图像中提取并恢复语义,降低左右图像光度对齐语义信息的损失,提升深度推断精度;构建2D语义提取模块,识别并提取目标周围的语义信息,强化关键区域的语义传输;最后通过融合网络融合恢复的语义,重构立体视觉图像以支持3D检测。仿真结果表明,所提方法使检测精度提升近70%,在低信噪比环境下显著优于传统方法。
原文摘要 · Abstract (English)
With the development of computer vision, 3D object detection has become increasingly important in many real-world applications. Limited by the computing power of sensor-side hardware, the detection task is sometimes deployed on remote computing devices or the cloud to execute complex algorithms, which brings massive data transmission overhead. In response, this paper proposes an optical flow-driven semantic communication framework for the stereo-vision 3D object detection task. The proposed framework fully exploits the dependence of stereo-vision 3D detection on semantic information in images and prioritizes the transmission of this semantic information to reduce total transmission data sizes while ensuring the detection accuracy. Specifically, we develop an optical flow-driven module to jointly extract and recover semantics from the left and right images to reduce the loss of the left-right photometric alignment semantic information and improve the accuracy of depth inference. Then, we design a 2D semantic extraction module to identify and extract semantic meaning around the objects to enhance the transmission of semantic information in the key areas. Finally, a fusion network is used to fuse the recovered semantics, and reconstruct the stereo-vision images for 3D detection. Simulation results show that the proposed method improves the detection accuracy by nearly 70% and outperforms the traditional method, especially for the low signal-to-noise ratio regime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。