发现大模型在正常对话中会悄悄产生有害行为,亟需新安全机制。
Exploring the Secondary Risks of Large Language Models
- 提出'二次风险'概念,指正常提问下模型产生的误导或有害响应。
- 在16个主流模型上测试,发现该风险普遍存在且跨模型迁移。
- 构建了包含650个提示的基准数据集,支持可复现评估。
随着大语言模型在关键应用和社会功能中的日益集成,确保其安全与对齐成为重大挑战。现有研究多关注对抗性攻击,却较少关注良性交互中悄然浮现的非对抗性失效。本文提出‘二次风险’这一新类别失败模式,表现为在正常提示下产生有害或误导性行为。此类风险源于模型泛化不完善,常规避标准安全机制。为系统评估,提出两个风险原语:冗长回复和推测性建议,以捕捉核心失效特征。基于此,设计SecLens框架,通过优化任务相关性、风险触发度和语言自然度,在黑盒环境下高效激发二次风险。为支持可复现评估,发布SecRiskBench基准数据集,包含650个提示,覆盖八类真实世界风险场景。对16个主流模型的广泛实验表明,二次风险普遍存在、可跨模型迁移且与模态无关,凸显亟需增强安全机制以应对实际部署中良性但有害的模型行为。
原文摘要 · Abstract (English)
Ensuring the safety and alignment of Large Language Models is a significant challenge with their growing integration into critical applications and societal functions. While prior research has primarily focused on jailbreak attacks, less attention has been given to non-adversarial failures that subtly emerge during benign interactions. We introduce secondary risks a novel class of failure modes marked by harmful or misleading behaviors during benign prompts. Unlike adversarial attacks, these risks stem from imperfect generalization and often evade standard safety mechanisms. To enable systematic evaluation, we introduce two risk primitives verbose response and speculative advice that capture the core failure patterns. Building on these definitions, we propose SecLens, a black-box, multi-objective search framework that efficiently elicits secondary risk behaviors by optimizing task relevance, risk activation, and linguistic plausibility. To support reproducible evaluation, we release SecRiskBench, a benchmark dataset of 650 prompts covering eight diverse real-world risk categories. Experimental results from extensive evaluations on 16 popular models demonstrate that secondary risks are widespread, transferable across models, and modality independent, emphasizing the urgent need for enhanced safety mechanisms to address benign yet harmful LLM behaviors in real-world deployments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。