用注意力机制筛选高质量长文本数据,仅用10亿词元就提升大模型长文本处理能力
LADM: Long-context Training Data Selection with Attention-based Dependency Measurement for LLMs
- 基于注意力机制捕捉上下文依赖,评估长文本数据质量
- 仅用10亿词元持续训练,显著提升多个长文本任务表现
- 适合需要优化长文本理解的大模型研究与应用
长上下文建模在大语言模型领域受到越来越多关注。通过持续使用长上下文数据进行训练,已成为赋予大模型处理长输入能力的标准方法。然而,如何衡量长上下文训练数据的质量仍是一个开放挑战。为此,我们提出一种基于注意力依赖度量的长文本数据选择框架(LADM),能够从大规模多领域预训练语料库中高效识别高质量长文本数据。LADM利用注意力机制的检索能力捕捉上下文依赖关系,确保对长文本数据质量的全面评估。实验结果表明,仅需10亿词元的持续训练,该框架即可显著提升大模型在多个长文本任务上的性能。
原文摘要 · Abstract (English)
Long-context modeling has drawn more and more attention in the area of Large Language Models (LLMs). Continual training with long-context data becomes the de-facto method to equip LLMs with the ability to process long inputs. However, it still remains an open challenge to measure the quality of long-context training data. To address this issue, we propose a Long-context data selection framework with Attention-based Dependency Measurement (LADM), which can efficiently identify high-quality long-context data from a large-scale, multi-domain pre-training corpus. LADM leverages the retrieval capabilities of the attention mechanism to capture contextual dependencies, ensuring a comprehensive quality measurement of long-context data. Experimental results show that our LADM framework significantly boosts the performance of LLMs on multiple long-context tasks with only 1B tokens for continual training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。