arXiv:2510.23198cs.LGcs.AI2025-10被引 1

提出可预测未见训练预算下领域适应性能的新规律

PTPP-Aware Adaptation Scaling Laws: Predicting Domain-Adaptation Performance at Unseen Pre-Training Budgets

  • 将预训练预算作为显式变量,建立与领域适应性能的映射关系
  • 在多语言迁移中,早期模型可准确预测高预算下的目标损失
  • 帮助在算力限制下规划重放比例与训练预算,兼顾性能与遗忘控制

持续预训练(CPT)在领域适应中需平衡目标域收益与基础域稳定性。现有CPT尺度律通常假设固定预训练预算,难以预测不同参数-令牌比(PTPP)下的适应效果。本文提出面向PTPP的适应尺度律,将预训练预算作为显式变量,实现对未见过的PTPP下适应损失的准确预测。在英/阿→法多语言设置中,基于早期阶段(PTPP={15,31})训练的模型,能有效预测PTPP=279时的目标损失,且在Huber-on-log、MAE$_\mathrm{rel}$、校准斜率等指标上优于传统的无PTPP感知的DCPT迁移基线;完整诊断结果(RMSE、MAPE)见附录。此外,该方法还可用于实际场景:在计算资源受限条件下,规划满足目标性能与遗忘约束的重放比例与适配令牌预算。

原文摘要 · Abstract (English)

Continual pre-training (CPT) for domain adaptation must balance target-domain gains with stability on the base domain. Existing CPT scaling laws typically assume a fixed pre-training budget, which limits their ability to forecast adaptation outcomes for models trained at different tokens-per-parameter (PTPP). We present \emph{PTPP-aware} adaptation scaling laws that make the pre-training budget an explicit variable, enabling accurate \emph{prediction} of adaptation loss at unseen \ptpp. On a multilingual setup (English/Arabic $\rightarrow$ French), PTPP-aware formulations trained on early stages (\ptpp{}=\{15,31\}) predict target loss at \ptpp{}=279 and outperform a PTPP-agnostic \dcpt{} transfer baseline on metrics (Huber-on-log, MAE$_\mathrm{rel}$, calibration slope); full diagnostics (RMSE, MAPE) are in the appendix. Beyond forecasting, we show a practical use case: planning replay ratios and adaptation token budgets that satisfy target and forgetting constraints under compute limits.

领域适应尺度律持续学习多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。