用深度学习实现图像与点云的高效压缩,提升编码效率。
Learned Compression for Images and Point Clouds
- 动态自适应熵模型,通过传输编码分布本身实现高效压缩。
- 专用于分类任务的轻量点云编码器,比特率显著降低。
- 分析视频帧间运动在卷积隐空间中的表现形式。
过去十年中,深度学习在图像分类、超分辨率和风格迁移等计算机视觉任务中取得了巨大成功。如今,我们将其应用于数据压缩,以推动下一代多媒体编解码器的发展。本论文在这一新兴领域做出三项主要贡献:首先,提出一种高效的低复杂度熵模型,通过将编码分布本身作为辅助信息进行压缩和传输,动态适应特定输入;其次,提出一种新型轻量级低复杂度点云编解码器,高度专用于分类任务,在比特率上显著优于非专用编解码器;最后,探索了连续视频帧间运动在卷积生成的隐空间中的表现形式。
原文摘要 · Abstract (English)
Over the last decade, deep learning has shown great success at performing computer vision tasks, including classification, super-resolution, and style transfer. Now, we apply it to data compression to help build the next generation of multimedia codecs. This thesis provides three primary contributions to this new field of learned compression. First, we present an efficient low-complexity entropy model that dynamically adapts the encoding distribution to a specific input by compressing and transmitting the encoding distribution itself as side information. Secondly, we propose a novel lightweight low-complexity point cloud codec that is highly specialized for classification, attaining significant reductions in bitrate compared to non-specialized codecs. Lastly, we explore how motion within the input domain between consecutive video frames is manifested in the corresponding convolutionally-derived latent space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。