JPEG推出点云编码新标准,让压缩后的数据既适合人看也适合机器用。
The JPEG Pleno Learning-based Point Cloud Coding Standard: Serving Man and Machine
- 用深度学习直接处理3D点云几何,2D投影编码颜色数据
- 相比传统方法,几何压缩率提升显著,颜色压缩略弱但整体高效
- 适合需要兼顾视觉与智能分析的虚拟现实、自动驾驶等场景
高效点云编码对虚拟现实、自动驾驶和数字孪生系统等应用日益关键,这些场景依赖丰富的交互式三维数据。深度学习在此领域展现强大能力,能比传统编码方法更高效地压缩点云,并支持在压缩域内直接进行计算机视觉任务,首次实现同时服务于人类观看和机器处理的统一压缩表示。为此,JPEG近日正式发布基于学习的点云编码(PCC)标准,可高效实现静态点云的有损编码,兼顾人眼感知与机器处理需求。几何部分通过稀疏卷积神经网络直接在原始3D形式上处理,颜色数据则投影至2D图像后采用基于学习的JPEG AI标准编码。本文全面描述该标准技术细节,并与当前最先进方法进行基准对比,揭示其主要优势与局限:在几何编码方面,相比传统MPEG PCC标准有显著率失真优势;颜色编码性能相对较弱,但得益于几何与颜色均采用端到端学习框架,以及高效的压缩域处理能力,整体表现仍具竞争力。
原文摘要 · Abstract (English)
Efficient point cloud coding has become increasingly critical for multiple applications such as virtual reality, autonomous driving, and digital twin systems, where rich and interactive 3D data representations may functionally make the difference. Deep learning has emerged as a powerful tool in this domain, offering advanced techniques for compressing point clouds more efficiently than conventional coding methods while also allowing effective computer vision tasks performed in the compressed domain thus, for the first time, making available a common compressed visual representation effective for both man and machine. Taking advantage of this potential, JPEG has recently finalized the JPEG Pleno Learning-based Point Cloud Coding (PCC) standard offering efficient lossy coding of static point clouds, targeting both human visualization and machine processing by leveraging deep learning models for geometry and color coding. The geometry is processed directly in its original 3D form using sparse convolutional neural networks, while the color data is projected onto 2D images and encoded using the also learning-based JPEG AI standard. The goal of this paper is to provide a complete technical description of the JPEG PCC standard, along with a thorough benchmarking of its performance against the state-of-the-art, while highlighting its main strengths and weaknesses. In terms of compression performance, JPEG PCC outperforms the conventional MPEG PCC standards, especially in geometry coding, achieving significant rate reductions. Color compression performance is less competitive but this is overcome by the power of a full learning-based coding framework for both geometry and color and the associated effective compressed domain processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。