arXiv:2409.07325cs.ITcs.LG2024-09被引 2

提出一种统计可验证的瓶颈压缩方法,确保特征满足信息约束。

Statistically Valid Information Bottleneck via Multiple Hypothesis Testing

  • 基于多重假设检验构建新框架,无需调参即可保证约束
  • 在语言模型压缩任务中显著提升统计可靠性
  • 适合需要严格理论保障的机器学习应用

信息瓶颈(IB)是机器学习中提取对下游任务有信息量的压缩特征的广泛研究框架。然而,现有方法依赖启发式超参数调优,无法保证学习到的特征满足信息论约束。本文提出一种统计有效的解决方案——通过多重假设检验的IB(IB-MHT),无论数据集规模如何,均能以高概率确保学习特征满足IB约束。该方法基于帕累托检验和学-测(LTT)范式,可封装现有IB求解器,为IB约束提供统计保证。我们在经典与确定性IB形式下验证了IB-MHT性能,包括语言模型蒸馏实验。结果表明,IB-MHT在统计稳健性和可靠性方面优于传统方法。

原文摘要 · Abstract (English)

The information bottleneck (IB) problem is a widely studied framework in machine learning for extracting compressed features that are informative for downstream tasks. However, current approaches to solving the IB problem rely on a heuristic tuning of hyperparameters, offering no guarantees that the learned features satisfy information-theoretic constraints. In this work, we introduce a statistically valid solution to this problem, referred to as IB via multiple hypothesis testing (IB-MHT), which ensures that the learned features meet the IB constraints with high probability, regardless of the size of the available dataset. The proposed methodology builds on Pareto testing and learn-then-test (LTT), and it wraps around existing IB solvers to provide statistical guarantees on the IB constraints. We demonstrate the performance of IB-MHT on classical and deterministic IB formulations, including experiments on distillation of language models. The results validate the effectiveness of IB-MHT in outperforming conventional methods in terms of statistical robustness and reliability.

信息瓶颈统计保证模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。