用极低码率实现高质量视觉通信与分析,兼顾画面还原与任务精度。
Generative Semantic Coding for Ultra-Low Bitrate Visual Communication and Analysis
- 融合生成模型与图像压缩,用文本和编码隐向量协同引导精确重建。
- 在极低码率下达到与现有方法相当的图像质量与视觉分析准确率。
- 适合深空探测、战场侦察等极端带宽场景,尤其关注远程交互与分析。
针对深空探测、战场情报、复杂环境机器人导航等极端低带宽场景下的远程视觉分析与人机交互需求,本文提出一种超低码率视觉通信新范式。传统文本生成模型仅能提供语义级近似,难以满足视觉通信与分析需求。为此,我们创新性地将图像生成与深度图像压缩无缝融合,利用联合文本与编码隐向量引导修正流模型,实现对视觉场景的精准重建。语义文本描述与编码隐向量均以极低码率编码并传输至解码端。实验表明,该方法在显著降低带宽消耗的同时,可保持与现有方法相当的图像重建质量与视觉分析精度。代码将在论文录用后公开。
原文摘要 · Abstract (English)
We consider the problem of ultra-low bit rate visual communication for remote vision analysis, human interactions and control in challenging scenarios with very low communication bandwidth, such as deep space exploration, battlefield intelligence, and robot navigation in complex environments. In this paper, we ask the following important question: can we accurately reconstruct the visual scene using only a very small portion of the bit rate in existing coding methods while not sacrificing the accuracy of vision analysis and performance of human interactions? Existing text-to-image generation models offer a new approach for ultra-low bitrate image description. However, they can only achieve a semantic-level approximation of the visual scene, which is far insufficient for the purpose of visual communication and remote vision analysis and human interactions. To address this important issue, we propose to seamlessly integrate image generation with deep image compression, using joint text and coding latent to guide the rectified flow models for precise generation of the visual scene. The semantic text description and coding latent are both encoded and transmitted to the decoder at a very small bit rate. Experimental results demonstrate that our method can achieve the same image reconstruction quality and vision analysis accuracy as existing methods while using much less bandwidth. The code will be released upon paper acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。