arXiv:2605.15507cs.ITcs.AI2026-05

提出新型向量量化方法,高效压缩混合高斯信号。

PrismQuant: Rate-Distortion-Optimal Vector Quantization for Gaussian-Mixture Sources

论文配图:PrismQuant: Rate-Distortion-Optimal Vector Quantization for Gaussian-Mixture Sources
图 1 · 摘自论文原文
  • 基于混合高斯源结构,分离成分量标签与残差编码
  • 理论逼近率失真极限,实测性能优于现有模型
  • 适合小模型、高精度信号压缩场景

对于均方误差下的高斯源,经典变换编码在率失真(RD)上是最优的:卡亨-洛埃变换(KLT)对角化协方差,逆水填法分配比特,标量量化闭环完成。这一优雅框架在多模态源中失效,因单一协方差无法捕捉异质局部几何,且率失真函数失去闭式表达。本文通过高斯混合源重新审视该问题,构建可构造的率失真理论。关键发现是混合结构仅引入分量标签开销;条件于活跃分量,各分支均为高斯分布,挑战在于跨异构分支的比特分配。我们证明,天使辅助的条件率失真函数由单一全局逆水填水平统一控制所有分量和特征模式。基于此,提出PrismQuant:无损传输分量标签,用匹配分量的KLT编码残差,再经标量量化,达到每维源分量的理论率界上界,渐近间隙趋零。进一步开发基于EM学习高斯混合、分量自适应KLT及熵约束标量量化(ECSQ)的实用实现。合成高斯混合实验表明,PrismQuant逼近理论率失真界限;真实信道状态信息(CSI)数据实验显示,相比基于Transformer的端到端编码器,在模型规模小一个数量级以上时仍具竞争力或更优性能。

原文摘要 · Abstract (English)

For a Gaussian source under mean-squared error (MSE), classical transform coding is rate--distortion (RD) optimal: the Karhunen--Loeve transform (KLT) diagonalizes the covariance, reverse waterfilling allocates the bits, and scalar quantization closes the loop. This elegant story breaks down for multimodal sources, where no single covariance can capture heterogeneous local geometries, and the RD function loses its closed form. We revisit this problem through Gaussian-mixture sources and develop a constructive RD theory for them. Our key finding is that the mixture structure incurs only a component label cost. Conditioned on the active mixture component, each branch is Gaussian; the challenge is allocating bits across heterogeneous branches. We prove that the genie-aided conditional RD function is governed by a single global reverse-waterfilling level shared across all components and eigenmodes. Building on this result, we introduce PrismQuant, which transmits the component label losslessly and encodes the residual using the component-matched KLT, followed by scalar quantization, achieving a rate of H(C)/n bits per source dimension of the converse, with a vanishing asymptotic gap. We further develop a practical implementation based on EM-driven Gaussian-mixture learning, component-adaptive KLTs, and entropy-constrained scalar quantization (ECSQ). Experiments on synthetic Gaussian mixtures show that PrismQuant closely approaches the theoretical RD bound, while experiments on real-world channel-state-information (CSI) data demonstrate competitive or superior performance compared with transformer-based learned codecs at more than one order of magnitude smaller model size.

向量量化率失真优化高斯混合小模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。