arXiv:2410.15633cs.CLcs.AI2024-10EMNLP被引 19

通过识别长依赖样本提升大模型长上下文理解能力

GATEAU: Selecting Influential Samples for Long Context Alignment

  • 基于长程依赖难度筛选关键训练样本
  • 选中的样本使模型长上下文理解能力显著提升
  • 适合研究长文本生成与指令遵循的学者

让大语言模型处理极长上下文指令的研究仍不充分。以往方法通过合成长指令跟随样本扩大数据量,但缺乏质量控制策略,易引入低质样本并限制模型性能。为此,我们提出GATEAU框架,通过评估两个关键维度——因长程依赖导致响应生成难度,以及因长程依赖导致输入理解难度——来识别富含长程依赖关系的关键样本。大量实验表明,使用这些精选样本训练的模型在指令遵循和长上下文理解方面表现更优。

原文摘要 · Abstract (English)

Aligning large language models to handle instructions with extremely long contexts has yet to be fully investigated. Previous studies have attempted to scale up the available data volume by synthesizing long instruction-following samples, as constructing such a dataset tends to be challenging for annotators. However, a lack of a well-defined strategy for ensuring data quality may introduce low-quality samples and restrict the model's performance. Thus, we propose GATEAU, a novel framework to address the unique challenge of long context alignment by identifying the influential samples enriched with long-range dependency relations. Specifically, GATEAU measures the long-range dependencies from two essential aspects: the difficulty of generating target responses due to the long-range dependencies, and the difficulty of understanding long inputs due to such dependencies. Comprehensive experiments indicate that GATEAU effectively identifies influential samples, and the model trained on these selected samples exhibits better instruction-following and long-context understanding capabilities.

长上下文样本选择大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。