arXiv:2605.07546cs.LG2026-05

发现数据变换中的不变性,让模型缩放规律跨领域通用。

On the Invariance and Generality of Neural Scaling Laws

  • 识别出保持信息的变换下缩放规律不变
  • 非双射变换降低信息分辨率ρ,规律变化可预测
  • 在文本、视觉、语音中验证,适用于医疗和时间序列任务

神经缩放定律建立了模型性能与数据量或计算资源之间的可预测关系,为新领域和任务的资源分配提供关键指导。然而,这些定律最需要时却最难获得:为新模型-任务对拟合缩放定律需昂贵的全面搜索,往往耗尽本应节省的算力预算。本文提出如何构建通用缩放定律:在资源充足的源域中拟合一次,即可可靠迁移到无法运行完整搜索的新域。这需要理解缩放特性何时及为何改变。我们通过识别正确的不变量解决该问题:在数据的双射(信息保持)变换下,缩放规律保持不变;在降低信息分辨率ρ的非双射变换下,规律以信息论为基础可预测地改变,形成单一可迁移轴。我们在语言、视觉和语音任务中验证该理论,并展示两个跨域应用:从通用文本拟合的规律预测电子健康记录上训练的语言模型缩放性能,以及在不同噪声水平下预测时间序列分类的缩放行为,恢复数据缩放指数误差小于3%。

原文摘要 · Abstract (English)

Neural scaling laws establish a predictable relationship between model performance and data or compute, offering crucial guidance for resource allocation in new domains and tasks. Yet such laws are most needed precisely where they are hardest to obtain: fitting one for a new model task pair demands expensive sweeps that typically exhaust the very compute budget the law is meant to economize. This paper poses the research question of how to develop generalizable scaling laws: laws fit once on a well-resourced source domain and reliably transported to new domains where running a full sweep is infeasible, which requires a fundamental understanding of when and why scaling properties change. We address this by identifying the right invariants: scaling laws are preserved under bijective (information-preserving) transformations of the data and modified in predictable, information-theoretically grounded ways under non-bijective transformations that lower its information resolution $ρ$: a single axis along which a law fit in one domain can be transported to another. We validate this across language, vision, and speech, and demonstrate two cross-domain applications: predicting scaling for language models trained on electronic health records from laws fit on general text, and predicting time-series classification scaling under varying levels of noise injection, recovering the data-scaling exponents to within $3\%$ error.

缩放定律跨域迁移信息论通用性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。