用几何约束与多智能体反思生成有害文本数据,提升防护模型效果
GRAID: Synthetic Data Generation with Geometric Constraints and Multi-Agentic Reflection for Harmful Content Detection
- 通过几何约束控制生成文本,确保语义覆盖全面
- 多智能体反思机制提升风格多样性,发现边缘案例
- 适用于需要增强有害内容检测能力的AI安全场景
针对守护类应用中有害文本分类的数据稀缺问题,我们提出GRAID(几何与反思驱动的AI数据增强)框架,利用大语言模型进行数据集扩充。该方法包含两个阶段:(i) 使用受限大模型生成几何可控的样本;(ii) 通过多智能体反思过程实现数据增强,提升风格多样性并挖掘边缘案例。该组合实现了输入空间的可靠覆盖与有害内容的细致探索。在两个基准数据集上验证表明,使用GRAID扩充有害文本分类数据集可显著提升下游防护模型性能。
原文摘要 · Abstract (English)
We address the problem of data scarcity in harmful text classification for guardrailing applications and introduce GRAID (Geometric and Reflective AI-Driven Data Augmentation), a novel pipeline that leverages Large Language Models (LLMs) for dataset augmentation. GRAID consists of two stages: (i) generation of geometrically controlled examples using a constrained LLM, and (ii) augmentation through a multi-agentic reflective process that promotes stylistic diversity and uncovers edge cases. This combination enables both reliable coverage of the input space and nuanced exploration of harmful content. Using two benchmark data sets, we demonstrate that augmenting a harmful text classification dataset with GRAID leads to significant improvements in downstream guardrail model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。