用测验提升训练效率,让大模型更高效学领域知识
TELLME: Test-Enhanced Learning for Language Model Enrichment

- 训练时插入测验题,利用反馈优化学习过程
- 金融领域性能提升23.6%,长期记忆保留率提高9.8%
- 适合需要高效适应新领域且资源有限的场景
持续预训练(CPT)被广泛用于大语言模型的领域适配,但面临获取大规模领域数据困难和计算成本高的挑战。本文提出测试增强学习方法TELLME,将测试增强学习(TEL)原理与CPT结合,通过在训练中引入测验题提升模型训练效率,促进领域知识高效获取和长期记忆保留。实验表明,TELLME在金融领域性能较现有方法最高提升23.6%,长期记忆保留率提升9.8%。
原文摘要 · Abstract (English)
Continual pre-training (CPT) has been widely adopted as a method for domain adaptation in large language models. However, CPT has consistently been accompanied by challenges, such as the difficulty of acquiring large-scale domain-specific datasets and high computational costs. In this study, we propose a novel method called Test-Enhanced Learning for Language Model Enrichment (TELLME) to alleviate these issues. TELLME leverages the TestEnhanced Learning (TEL) principle, whereby the model's training efficiency is improved using quizzes during training. It integrates this principle with CPT, thereby promoting efficient domain-specific knowledge acquisition and long-term memory retention. Experimental results demonstrate that TELLME outperforms existing methods by up to 23.6% in the financial domain and achieves a 9.8% improvement in long-term memory retention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。