通过神经元激活模式筛选高质量指令数据,提升大模型能力。
Neuron-Aware Data Selection In Instruction Tuning For Large Language Models
- 基于神经元激活相似性筛选指令数据,无需依赖外部模型。
- 仅用10%的Alpaca-GPT4数据即超越现有方法在多任务表现。
- 逻辑与编程类数据具备强迁移能力,少量核心数据可通用提升性能。
指令微调(IT)已被证明是激发大语言模型(LLMs)强大能力的有效方法。近期研究指出,过多的IT数据会降低模型性能,而精心挑选的小规模高质量数据子集能显著增强其能力。因此,从IT数据集中高效识别出能有效发展特定或通用能力的数据子集成为关键挑战。为此,我们提出一种新颖高效的框架NAIT。NAIT通过分析IT数据与目标领域能力之间的神经元激活模式相似性来评估其影响。具体而言,NAIT从目标领域内的域内数据中捕捉神经元激活模式,构建可复用且可迁移的神经元激活特征;随后根据候选样本与目标能力预期激活特征的相似性进行评估与选择。实验结果表明,使用NAIT筛选出的10% Alpaca-GPT4 IT数据子集,在多个任务上均持续优于依赖外部先进模型或基于不确定性的方法。研究还发现神经元激活特征在不同能力间具有可迁移性:具有更多逻辑推理和程序化特征的IT数据具备强泛化迁移能力,能够促进模型在多任务上的综合提升;而一个稳定的最小数据子集足以持续激活模型基础能力,并普遍提升多样任务的表现。
原文摘要 · Abstract (English)
Instruction Tuning (IT) has been proven to be an effective approach to unlock the powerful capabilities of large language models (LLMs). Recent studies indicate that excessive IT data can degrade LLMs performance, while carefully selecting a small subset of high-quality IT data can significantly enhance their capabilities. Therefore, identifying the most efficient subset data from the IT dataset to effectively develop either specific or general abilities in LLMs has become a critical challenge. To address this, we propose a novel and efficient framework called NAIT. NAIT evaluates the impact of IT data on LLMs performance by analyzing the similarity of neuron activation patterns between the IT dataset and the target domain capability. Specifically, NAIT captures neuron activation patterns from in-domain datasets of target domain capabilities to construct reusable and transferable neuron activation features. It then evaluates and selects optimal samples based on the similarity between candidate samples and the expected activation features of the target capabilities. Experimental results show that training on the 10\% Alpaca-GPT4 IT data subset selected by NAIT consistently outperforms methods that rely on external advanced models or uncertainty-based features across various tasks. Our findings also reveal the transferability of neuron activation features across different capabilities of LLMs. In particular, IT data with more logical reasoning and programmatic features possesses strong general transferability, enabling models to develop stronger capabilities across multiple tasks, while a stable core subset of data is sufficient to consistently activate fundamental model capabilities and universally improve performance across diverse tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。