用少量财务数据就能让大模型高效专精,且不会忘记通用知识。
The Data Efficiency Frontier of Financial Foundation Models: Scaling Laws from Continued Pretraining
- 在4亿字节财务文本上继续预训练10亿和30亿参数模型
- 前2亿字节提升最明显,之后收益递减,损失持续下降
- 小预算即可实现专业领域优化,适合金融AI开发者参考
领域自适应预训练(DAPT)为在不完全重训的情况下将大语言模型专门化于高价值领域提供了可行路径。我们对美国证监会文件上的持续预训练进行了早期规模定律分析,使用10亿和30亿参数的Llama-3.2模型,在包含4亿词元的金融语料库上进行训练,并在5000万、1亿、2亿和4亿词元处设置验证检查点。结果表明,两个模型在证监会领域验证损失上均持续改善,最大提升集中在前2亿词元内,后续收益递减。幂律拟合显示指数较浅,说明金融语言高度规则且可高效学习。通用领域验证损失在所有词元预算下基本不变,表明无显著漂移或灾难性遗忘。数据效率前沿进一步显示,两个模型在提升专业化程度的同时,混合领域退化可忽略不计。这些发现为金融基础模型的扩展提供了早期实证指导,表明以相对较小的词元预算即可实现有效领域适配,且更大模型规模(70亿至700亿参数)在预期数据需求下仍具可行性。
原文摘要 · Abstract (English)
Domain-adaptive pretraining (DAPT) offers a practical path to specializing large language models for high-value domains without full retraining. We conduct an early-stage scaling-law analysis of continued pretraining on U.S. SEC filings, training 1B and 3B-parameter Llama-3.2 models on a 400M-token financial corpus with validation checkpoints at 50M, 100M, 200M, and 400M tokens. Results show consistent improvements in SEC-domain validation loss for both models, with the largest gains occurring within the first 200M tokens and diminishing returns thereafter. Power-law fits reveal shallow exponents, indicating that financial language is highly regular and efficiently learnable under continued pretraining. General-domain validation loss remains effectively unchanged across all token budgets, suggesting minimal drift and no signs of catastrophic forgetting. A data-efficiency frontier further shows that both models move toward improved specialization with negligible mixed-domain degradation. Together, these findings provide early empirical guidance for scaling financial foundation models, suggesting that meaningful domain adaptation can be achieved with comparatively modest token budgets and that larger model scales (7B-70B) remain tractable under projected data requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。