用LLM辅助发现并修复扩散模型的失败模式。
LLM-Assisted Red Teaming of Diffusion Models through "Failures Are Fated, But Can Be Faded"
- 结合强化学习与LLM生成反馈,系统探测模型缺陷。
- 在扩散模型上验证方法有效,可重构更安全的输出分布。
- 适合模型安全审计与对齐优化的研究者使用。
在看似表现良好的大型深度神经网络中,仍存在精度、社会偏见及与人类价值观对齐等方面的若干失败现象。因此,在部署前刻画这些失败模式对于工程师调试或审计模型至关重要。然而,穷举所有可能导致模型失败的因素组合不切实际。本文改进了‘失败注定,但可消解’框架(arXiv:2406.07145)——一种用于探索和构建预训练生成模型失败景观的后处理方法——引入多种深度强化学习算法、筛选测试以及基于LLM的奖励函数与状态生成机制。在少量人工反馈协助下,我们展示了如何通过避开已发现的失败模式,重构出更理想的失败景观。我们在扩散模型上实证验证了该方法的有效性,并分析了各算法在识别失败模式中的优劣。
原文摘要 · Abstract (English)
In large deep neural networks that seem to perform surprisingly well on many tasks, we also observe a few failures related to accuracy, social biases, and alignment with human values, among others. Therefore, before deploying these models, it is crucial to characterize this failure landscape for engineers to debug or audit models. Nevertheless, it is infeasible to exhaustively test for all possible combinations of factors that could lead to a model's failure. In this paper, we improve the "Failures are fated, but can be faded" framework (arXiv:2406.07145)--a post-hoc method to explore and construct the failure landscape in pre-trained generative models--with a variety of deep reinforcement learning algorithms, screening tests, and LLM-based rewards and state generation. With the aid of limited human feedback, we then demonstrate how to restructure the failure landscape to be more desirable by moving away from the discovered failure modes. We empirically demonstrate the effectiveness of the proposed method on diffusion models. We also highlight the strengths and weaknesses of each algorithm in identifying failure modes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。