用主动学习让大模型生成更全面的安全场景,避免遗漏关键罕见案例。
Active Learning for Robust and Representative LLM Generation in Safety-Critical Scenarios
- 结合主动学习与聚类,引导大模型生成更具代表性的安全场景。
- 构建5400条潜在安全违规数据,提升模型对罕见场景的覆盖。
- 无需先验分布知识,适用于多种安全检测任务,效果显著。
在面向用户的应用系统中,确保各类场景下的鲁棒性安全措施至关重要。尽管大型语言模型(LLMs)可生成有价值的安全数据,但其常存在分布偏差,过度关注常见场景而忽略罕见但关键的情况,从而削弱安全协议的有效性。为此,我们提出一种新框架,将主动学习与聚类相结合,引导LLM生成更具代表性和鲁棒性的安全场景。通过迭代过程,利用LLM生成并由主动学习模型反馈,我们构建了一个包含5.4K条潜在安全违规的数据集。结果表明,该方法在不依赖底层数据分布先验的情况下,能生成更全面的安全场景;同时,所获数据不仅提升了主动学习模型的准确率和F1分数,也改善了其他未参与主动学习过程的模型性能,展现出广泛适用性。
原文摘要 · Abstract (English)
Ensuring robust safety measures across a wide range of scenarios is crucial for user-facing systems. While Large Language Models (LLMs) can generate valuable data for safety measures, they often exhibit distributional biases, focusing on common scenarios and neglecting rare but critical cases. This can undermine the effectiveness of safety protocols developed using such data. To address this, we propose a novel framework that integrates active learning with clustering to guide LLM generation, enhancing their representativeness and robustness in safety scenarios. We demonstrate the effectiveness of our approach by constructing a dataset of 5.4K potential safety violations through an iterative process involving LLM generation and an active learner model's feedback. Our results show that the proposed framework produces a more representative set of safety scenarios without requiring prior knowledge of the underlying data distribution. Additionally, data acquired through our method improves the accuracy and F1 score of both the active learner model as well models outside the scope of active learning process, highlighting its broad applicability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。