arXiv:2505.23032cs.LGcs.AI2025-05ICML

用贝叶斯方法精准预测模型扩展性能,还能给出可信度。

Bayesian Neural Scaling Law Extrapolation with Prior-Data Fitted Networks

  • 基于先验-数据拟合网络构建可采样的合成函数分布
  • 在数据稀缺时仍保持高精度,优于传统点估计与贝叶斯方法
  • 适合需评估风险的决策场景,如资源投入规划

扩展是深度学习近年进步的主要驱动力。大量实证研究发现,扩展规律常呈幂律形式,并提出了多种幂律变体以预测更大规模下的表现。然而,现有方法多依赖点估计且不量化不确定性,而这一信息对涉及决策的问题(如评估额外计算资源带来的性能提升)至关重要。本文提出基于先验-数据拟合网络(PFNs)的贝叶斯框架,用于神经网络扩展规律外推。我们设计了一种先验分布,可生成无限多类比真实扩展规律的合成函数,使PFN能元学习外推能力。我们在真实神经网络扩展规律上验证了该方法的有效性,对比了现有点估计方法与贝叶斯方法。结果表明,本方法在数据受限场景(如贝叶斯主动学习)中表现更优,展现出在实际应用中实现可靠、带不确定性的外推潜力。

原文摘要 · Abstract (English)

Scaling has been a major driver of recent advancements in deep learning. Numerous empirical studies have found that scaling laws often follow the power-law and proposed several variants of power-law functions to predict the scaling behavior at larger scales. However, existing methods mostly rely on point estimation and do not quantify uncertainty, which is crucial for real-world applications involving decision-making problems such as determining the expected performance improvements achievable by investing additional computational resources. In this work, we explore a Bayesian framework based on Prior-data Fitted Networks (PFNs) for neural scaling law extrapolation. Specifically, we design a prior distribution that enables the sampling of infinitely many synthetic functions resembling real-world neural scaling laws, allowing our PFN to meta-learn the extrapolation. We validate the effectiveness of our approach on real-world neural scaling laws, comparing it against both the existing point estimation methods and Bayesian approaches. Our method demonstrates superior performance, particularly in data-limited scenarios such as Bayesian active learning, underscoring its potential for reliable, uncertainty-aware extrapolation in practical applications.

贝叶斯推理模型扩展不确定性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。