用激活调控生成安全检测数据,发现多样性是关键
Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection

- 通过激活调控生成目标概念响应,引入样本与集合级多样性作为评估新维度
- 41/136种配置优于提示生成数据,多样性提升使下游AUROC平均提高12%
- 适合想优化合成数据质量的安全检测研究者,尤其关注参数调优策略
安全检测模型需要大量违反HHH(有益、无害、诚实)原则的输出样本以实现稳健泛化,但此类样本稀缺。激活调控(AS)作为一种数据高效的方法,可生成与目标概念对齐的响应。本文在4个概念×2个模型×4种调控方法下开展内外部评估。内在评估中,除标准的对齐度与连贯性外,首次引入样本级和集合级多样性作为质量维度,发现增强调控强度会降低响应多样性。外在评估中,将现有训练数据中的违规样本替换为调控生成结果,并微调检测分类器。结果显示,在4个概念中有3个上AS生成数据优于提示生成数据;但仅有41/136种配置表现更优,表明下游性能取决于对齐、连贯与多样性的协同满足。三者调和均值比单独使用对齐与连贯性更能一致预测下游AUROC,为实践者提供了可操作的超参调优参考。结果表明AS在合成数据生成中具有潜力,而多样性是此前被忽视的关键调节维度。
原文摘要 · Abstract (English)
Safety detection models require examples of HHH (Helpful, Harmless, Honest)-violating outputs for robust generalization, however such examples are scarce. Activation Steering (AS) has emerged as a data-efficient method for generating target-concept-aligned responses. We investigate whether AS can generate high-quality training datasets for downstream classifiers, a question that remains untested. We present a two-fold study with intrinsic and extrinsic evaluation across $4$ concepts $\times\,2$ models $\times\,4$ steering methods. Intrinsically, beyond the field-standard rubric of steering success (concept alignment) and coherence, we introduce sample- and set-level diversity as a quality axis previously absent from the literature, and find that increasing steering strength reduces response diversity. Extrinsically, we replace HHH-violating examples in the available training data with steered generations and fine-tune detection classifiers. AS-generated data results in a better classifier than the prompting-generated data on $3$ of $4$ concepts. However, only $41$ of $136$ AS configurations outperform prompting, indicating that downstream utility lies in a narrow regime that jointly satisfies success, coherence, and diversity. The harmonic mean of these three axes correlates with downstream AUROC more consistently across concepts than success and coherence alone, providing a practical heuristic target for practitioners tuning AS hyperparameters. Together, our results highlight the potential of AS in synthetic data generation for improving safety detection and identify diversity as a critical, previously overlooked axis for tuning AS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。