用动态反馈优化大模型安全对齐数据生成质量。
Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
- 基于预设原则动态监控数据生成,识别短板并指导迭代。
- 在三个主流大模型上提升安全对齐效果,不损失模型实用性。
- 适合需要高质量安全训练数据的研究者和开发者。
数据是大型语言模型(LLM)对齐的关键要素。近期研究探索了使用大模型进行高效数据收集,但生成数据常存在质量缺陷,如关键方面缺失或低质样本过多。为此,我们提出Data Advisor——一种增强型大模型驱动的数据生成方法,考虑目标数据集的特征。该方法从一组预定义原则出发,实时监测生成数据状态,识别当前数据集的薄弱环节,并据此指导下一阶段数据生成。Data Advisor可无缝集成至现有数据生成流程,显著提升数据质量与覆盖度。在Mistral、Llama2和Falcon三个代表性大模型上的安全对齐实验表明,该方法有效提升了模型对细粒度安全问题的抵御能力,同时未牺牲模型可用性。
原文摘要 · Abstract (English)
Data is a crucial element in large language model (LLM) alignment. Recent studies have explored using LLMs for efficient data collection. However, LLM-generated data often suffers from quality issues, with underrepresented or absent aspects and low-quality datapoints. To address these problems, we propose Data Advisor, an enhanced LLM-based method for generating data that takes into account the characteristics of the desired dataset. Starting from a set of pre-defined principles in hand, Data Advisor monitors the status of the generated data, identifies weaknesses in the current dataset, and advises the next iteration of data generation accordingly. Data Advisor can be easily integrated into existing data generation methods to enhance data quality and coverage. Experiments on safety alignment of three representative LLMs (i.e., Mistral, Llama2, and Falcon) demonstrate the effectiveness of Data Advisor in enhancing model safety against various fine-grained safety issues without sacrificing model utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。