用语义等价集重构信号,统一感知压缩的理论基础。
A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff

- 将重建目标从原始样本改为语义等价集合中的任意样本
- 提出相似性变分下界,实现可计算的语义压缩分析
- 揭示感知优化与语义信息提取的理论一致性,适合图像视频压缩研究者
自然信号压缩的根本极限传统上由经典率失真(RD)理论刻画,即码率与重建失真之间的权衡。而率失真感知(RDP)框架引入基于分布发散的感知质量度量作为建模原则,但其理论根源尚不明确。本文受语义等价性视角启发,将感知重建重新定义为恢复与源信号相关联的理想语义等价集(synset)中任一合法样本,而非源样本本身,并构建了语义等价源编码架构。在此基础上,提出语义等价变分推断(SVI)分析框架及语义等价变分下界(SVLBO),以实现对面向语义集压缩的可计算分析。该框架下建立了语义一致性原则,证明最优语义信息识别在理论上与感知优化一致。进一步导出紧致的语义等价源编码率表征,并表明其Jensen极限松弛形式即为适用于实际优化的语义率失真感知表达式。这些结果表明,分布发散项自然源自基于语义集的重建目标,阐明其与现有RDP形式及经典RD理论的兼容性,并暗示语义等价编码的潜在优势。
原文摘要 · Abstract (English)
The fundamental limit of natural signal compression has traditionally been characterized by classical rate-distortion (RD) theory through the tradeoff between coding rate and reconstruction distortion, while the rate-distortion-perception (RDP) framework introduces a divergence-based measure of perceptual quality as a modeling principle, leaving its theoretical origin unclear. In this paper, motivated by a synonymity-based semantic information perspective, we reformulate perceptual reconstruction as recovering any admissible sample within an ideal synonymous set (synset) associated with the source, rather than the source sample itself, and establish a synonymous source coding architecture. On this basis, we develop a synonymous variational inference (SVI) analysis framework with a synonymous variational lower bound (SVLBO) for tractable analysis of synset-oriented compression. Within this framework, we establish a synonymity-perception consistency principle, showing that optimal identification of semantic information is theoretically consistent with perceptual optimization. Based on this result, we further derive a tight-bound synonymous source coding rate characterization and show that its Jensen-limit relaxation leads to a synonymous rate-distortion-perception form for practical optimization. These analytical results show that the distributional divergence term arises naturally from the synset-based reconstruction objective, clarify its compatibility with existing RDP formulations and classical RD theory, and suggest the potential advantages of synonymous source coding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。