arXiv:2503.24028cs.AI2025-03

提升提示鲁棒性,筛选更高质量的指令数据

Pay More Attention to the Robustness of Prompt for Instruction Data Mining

  • 通过对抗攻击生成恶意指令,检验提示稳定性
  • 新指标衡量对抗指令对回复生成的帮助程度
  • 基于输出嵌入一致性筛选优质数据,适合大模型训练

指令微调已成为塑造大模型行为的关键方法。近期研究发现,仅需少量高质量指令数据即可实现良好性能。本文进一步探索提示鲁棒性对高质量指令数据筛选的影响,提出一种面向指令微调的在线高质量指令数据挖掘框架。核心创新在于:通过攻击在线指令数据的提示生成对抗指令数据;引入对抗指令遵循难度度量,评估对抗指令对生成响应的帮助程度;并提出对抗指令输出嵌入一致性方法,用于筛选高质量数据。在两个基准数据集上进行广泛实验,结果验证了两种方法的有效性,并凸显考虑提示鲁棒性的关键实践意义。

原文摘要 · Abstract (English)

Instruction tuning has emerged as a paramount method for tailoring the behaviors of LLMs. Recent work has unveiled the potential for LLMs to achieve high performance through fine-tuning with a limited quantity of high-quality instruction data. Building upon this approach, we further explore the impact of prompt's robustness on the selection of high-quality instruction data. This paper proposes a pioneering framework of high-quality online instruction data mining for instruction tuning, focusing on the impact of prompt's robustness on the data mining process. Our notable innovation, is to generate the adversarial instruction data by conducting the attack for the prompt of online instruction data. Then, we introduce an Adversarial Instruction-Following Difficulty metric to measure how much help the adversarial instruction data can provide to the generation of the corresponding response. Apart from it, we propose a novel Adversarial Instruction Output Embedding Consistency approach to select high-quality online instruction data. We conduct extensive experiments on two benchmark datasets to assess the performance. The experimental results serve to underscore the effectiveness of our proposed two methods. Moreover, the results underscore the critical practical significance of considering prompt's robustness.

指令微调提示鲁棒性数据挖掘大模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。