arXiv:2504.14240cs.CVcs.MM2025-04被引 22

基于兴趣区域的点云压缩,兼顾人眼与机器视觉质量。

ROI-Guided Point Cloud Geometry Compression Towards Human and Machine Vision

  • 双分支结构:基础层简化编码,增强层聚焦几何细节修复。
  • 引入ROI预测网络,提升高比特率下目标检测准确率10%。
  • 将注意力掩码融入码率-失真优化,适配机器感知任务。

点云数据在自动驾驶、虚拟现实和机器人等领域至关重要,但其庞大体积带来存储与传输挑战。现有压缩方法为追求高压缩比,常导致关键语义信息严重损失,影响下游任务精度。为此,我们首次提出面向人眼与机器视觉的区域兴趣(ROI)引导点云几何压缩(RPCGC)方法。框架采用双分支并行结构:基础层对简化点云进行编码解码,增强层则聚焦几何细节修复;同时,增强层残差经由ROI预测网络生成掩码,作为强监督信号注入残差,并精细应用于码率-失真(RD)优化中,每个点在失真计算中按掩码加权。损失函数包含RD损失与检测损失,以更好指导编码过程。实验表明,RPCGC在ScanNet与SUN RGB-D数据集上,于高比特率下实现优异压缩性能,检测准确率相比部分学习型压缩方法提升10%。

原文摘要 · Abstract (English)

Point cloud data is pivotal in applications like autonomous driving, virtual reality, and robotics. However, its substantial volume poses significant challenges in storage and transmission. In order to obtain a high compression ratio, crucial semantic details usually confront severe damage, leading to difficulties in guaranteeing the accuracy of downstream tasks. To tackle this problem, we are the first to introduce a novel Region of Interest (ROI)-guided Point Cloud Geometry Compression (RPCGC) method for human and machine vision. Our framework employs a dual-branch parallel structure, where the base layer encodes and decodes a simplified version of the point cloud, and the enhancement layer refines this by focusing on geometry details. Furthermore, the residual information of the enhancement layer undergoes refinement through an ROI prediction network. This network generates mask information, which is then incorporated into the residuals, serving as a strong supervision signal. Additionally, we intricately apply these mask details in the Rate-Distortion (RD) optimization process, with each point weighted in the distortion calculation. Our loss function includes RD loss and detection loss to better guide point cloud encoding for the machine. Experiment results demonstrate that RPCGC achieves exceptional compression performance and better detection accuracy (10% gain) than some learning-based compression methods at high bitrates in ScanNet and SUN RGB-D datasets.

点云压缩机器视觉注意力机制三维感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。