arXiv:2512.22638stat.MLcs.LG2025-12被引 1

让嵌入表示保留似然信息,实现压缩数据下的精准统计推断。

Likelihood-Preserving Embeddings for Statistical Inference

  • 提出似然比失真度量Δ_n,控制它即可保证推断不变。
  • 证明Δ_n = o_p(1)时,检验与估计结果与原始数据一致。
  • 用神经网络构造近似充分统计量,可从训练损失推导推断保证。

现代机器学习嵌入虽能高效压缩高维数据,却常破坏经典似然推断所需的几何结构。本文建立似然保持嵌入的严格理论:学习到的表示可替代原始数据用于假设检验、置信区间构建和模型选择,且不改变推断结论。提出似然比失真度量Δ_n,衡量嵌入引起的对数似然比最大误差。核心贡献为铰链定理:控制Δ_n是保持推断的充要条件。若Δ_n = o_p(1),则所有基于似然比的检验与贝叶斯因子渐近不变,且代理最大似然估计量渐近等价于全数据MLE。证明了通用似然保持需几乎可逆嵌入,推动特定模型类的保障设计。进而提出基于神经网络的构造框架,建立训练损失与推断保证间的显式边界。在高斯与柯西分布上的实验验证了指数族理论预测的尖锐相变,分布式临床推断应用展示了实用性。

原文摘要 · Abstract (English)

Modern machine learning embeddings provide powerful compression of high-dimensional data, yet they typically destroy the geometric structure required for classical likelihood-based statistical inference. This paper develops a rigorous theory of likelihood-preserving embeddings: learned representations that can replace raw data in likelihood-based workflows -- hypothesis testing, confidence interval construction, model selection -- without altering inferential conclusions. We introduce the Likelihood-Ratio Distortion metric $Δ_n$, which measures the maximum error in log-likelihood ratios induced by an embedding. Our main theoretical contribution is the Hinge Theorem, which establishes that controlling $Δ_n$ is necessary and sufficient for preserving inference. Specifically, if the distortion satisfies $Δ_n = o_p(1)$, then (i) all likelihood-ratio based tests and Bayes factors are asymptotically preserved, and (ii) surrogate maximum likelihood estimators are asymptotically equivalent to full-data MLEs. We prove an impossibility result showing that universal likelihood preservation requires essentially invertible embeddings, motivating the need for model-class-specific guarantees. We then provide a constructive framework using neural networks as approximate sufficient statistics, deriving explicit bounds connecting training loss to inferential guarantees. Experiments on Gaussian and Cauchy distributions validate the sharp phase transition predicted by exponential family theory, and applications to distributed clinical inference demonstrate practical utility.

统计推断嵌入表示似然保持神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。