在计算与速率约束下,用后验设计提升语义压缩效率。
Semantic Rate Distortion and Posterior Design: Compute Constraints, Multimodality, and Strategic Inference
- 通过后验协方差设计,在资源受限下优化语义压缩。
- 模型深度和推理时间每翻倍,语义准确率呈指数提升。
- 解释多模态大模型为资源约束下的后验设计机制。
研究在速率与计算约束下,编码器与解码器各自优化二次目标的策略性高斯语义压缩。潜在高斯状态生成任务相关的语义变量,解码器通过最小均方误差估计最优响应,使编码器问题转化为信息速率约束下的后验协方差设计。我们刻画了直接、远程和全信息情形下的战略率失真函数,推导出语义水填与速率约束高斯说服解,并证明在目标不一致时高斯解仍最优。进一步表明,架构计算限制等同于隐式速率约束,模型深度与推理时间计算每翻倍,语义准确率实现指数级提升;多模态观测可消除远程编码固有的几何均值惩罚。这些结果为数据与能源高效人工智能提供信息论基础,并为现代多模态语言模型提供了资源约束下的后验设计解释。
原文摘要 · Abstract (English)
We study strategic Gaussian semantic compression under rate and compute constraints, where an encoder and decoder optimize distinct quadratic objectives. A latent Gaussian state generates a task dependent semantic variable, and the decoder best responds via MMSE estimation, reducing the encoder's problem to posterior covariance design under an information rate constraint. We characterize the strategic rate distortion function in direct, remote, and full information regimes, derive semantic waterfilling and rate constrained Gaussian persuasion solutions, and establish Gaussian optimality under misaligned objectives. We further show that architectural compute limits act as implicit rate constraints, yielding exponential improvements in semantic accuracy with model depth and inference time compute, while multimodal observation eliminates the geometric mean penalty inherent to remote encoding. These results provide information theoretic foundations for data and energy efficient AI and offer a principled interpretation of modern multimodal language models as posterior design mechanisms under resource constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。