arXiv:2510.11842cs.LGcs.CL2025-10

通过实验找到合成数据与回放的最佳平衡,降低模型训练成本。

Balancing Synthetic Data and Replay for Enhancing Task-Specific Capabilities

  • 用合成数据+回放策略,在有限计算资源下适应新任务
  • 发现特定回放比例能兼顾新任务表现与旧知识保留
  • 为不同预算提供可落地的训练配置建议

通过持续预训练将语言模型适配新任务面临根本性权衡:需学习新能力的同时避免灾难性遗忘。尽管已有研究关注合成数据生成,但在计算资源受限下,平衡任务性能与知识保留的最优回放比例仍不明确。本文以bAbI推理任务为目标,系统评估不同总令牌预算和回放比例配置的影响。实验揭示了在任务掌握与通用知识保留间达到平衡的最优配置。基于结果,我们提出了基于计算预算的回放比例选择指南,使从业者能在显著降低训练成本的前提下实现强任务适应。

原文摘要 · Abstract (English)

Adapting language models to new tasks through continued pretraining faces a fundamental trade-off: models must learn new capabilities while avoiding catastrophic forgetting of existing knowledge. While prior work has studied synthetic data generation techniques, the optimal replay ratios for balancing task performance and knowledge retention under computational constraints remain poorly understood. We present a comprehensive empirical study investigating the interplay between replay ratio configuration and computational budget when adapting language models to new tasks. Using the bAbI reasoning tasks as our target objective, we apply synthetic data generation and systematically evaluate different total token budgets and replay ratio configurations. We analyze their effects on both task mastery and general knowledge retention. Our experiments reveal an optimal configuration that balances task-specific performance with general knowledge retention. Based on our findings, we provide empirically-grounded guidelines for selecting replay ratios based on computational budget, enabling practitioners to achieve strong task adaptation with significantly reduced training costs.

语言模型持续学习合成数据回放策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。