arXiv:2502.03843cs.CLcs.AI2025-02AAAI被引 6

用合成数据提升大模型的自然语言理解能力

Improving Natural Language Understanding for LLMs via Large-Scale Instruction Synthesis

  • 通过人机协作生成多样化的指令数据
  • 在5个NLU任务上平均提升3.1%性能
  • 适合想增强模型理解能力的研究者

高质量大规模指令对对齐大语言模型至关重要,但自然语言理解(NLU)领域存在严重指令短缺。现有工作多聚焦信息抽取(IE),忽视阅读理解、问答和文本分类等任务,且数据多样性不足导致模型泛化能力下降。为此,我们提出Hum——一个大规模、高质量的合成NLU指令语料库,涵盖信息抽取(闭式与开式)、阅读理解、文本分类及指令通用任务,丰富任务类型。同时引入人-大模型协同机制,通过指南、偏好规则和格式变体提升指令多样性。我们在5个NLU任务和28个通用能力评估数据集上进行实验,结果表明,Hum使6个大模型的NLU能力平均提升3.1%,且未出现其他通用能力显著下降。

原文摘要 · Abstract (English)

High-quality, large-scale instructions are crucial for aligning large language models (LLMs), however, there is a severe shortage of instruction in the field of natural language understanding (NLU). Previous works on constructing NLU instructions mainly focus on information extraction (IE), neglecting tasks such as machine reading comprehension, question answering, and text classification. Furthermore, the lack of diversity in the data has led to a decreased generalization ability of trained LLMs in other NLU tasks and a noticeable decline in the fundamental model's general capabilities. To address this issue, we propose Hum, a large-scale, high-quality synthetic instruction corpus for NLU tasks, designed to enhance the NLU capabilities of LLMs. Specifically, Hum includes IE (either close IE or open IE), machine reading comprehension, text classification, and instruction generalist tasks, thereby enriching task diversity. Additionally, we introduce a human-LLMs collaborative mechanism to synthesize instructions, which enriches instruction diversity by incorporating guidelines, preference rules, and format variants. We conduct extensive experiments on 5 NLU tasks and 28 general capability evaluation datasets for LLMs. Experimental results show that Hum enhances the NLU capabilities of six LLMs by an average of 3.1\%, with no significant decline observed in other general capabilities.

自然语言理解指令生成大模型训练合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。