arXiv:2504.19674cs.CRcs.AI2025-04EMNLP被引 10

SAGE通过模拟真实对话动态评估大模型安全,发现模型表现随对话轮次和用户性格显著变化。

SAGE: A Generic Framework for LLM Safety Evaluation

  • 用五大人格模型生成对抗性对话代理,实现多轮情境化测试
  • 长对话中危害率上升,同一模型在不同人格下表现差异大
  • 适合关注模型实际应用安全的开发者与评测团队

随着大语言模型在医疗、金融建议等场景快速部署,安全评估难以跟上步伐。现有基准仅关注单轮交互与通用规则,无法捕捉真实使用中的对话动态及上下文相关的特定风险,导致部分危害在标准评测中被忽略。为此,我们提出SAGE(Safety AI Generic Evaluation),一个自动化模块化框架,支持定制化、动态化的危害评估。SAGE基于五大人格模型生成具有多样性格的提示式对抗代理,实现系统感知的多轮对话,可适配目标应用场景与危害策略。我们在三个应用场景和多种危害策略下评估了七种先进大模型。多轮实验表明:危害随对话轮次增加而上升;模型行为在不同用户性格与场景下差异显著;部分模型通过高拒绝率降低危害,但严重影响实用性。此外,在同一危害类别内,收紧儿童性保护策略会显著提高各应用中的缺陷检测率。这些结果强调了面向政策、上下文和动态演化的安全测试对实际部署的重要性。

原文摘要 · Abstract (English)

As Large Language Models are rapidly deployed across diverse applications from healthcare to financial advice, safety evaluation struggles to keep pace. Current benchmarks focus on single-turn interactions with generic policies, failing to capture the conversational dynamics of real-world usage and the application-specific harms that emerge in context. Such potential oversights can lead to harms that go unnoticed in standard safety benchmarks and other current evaluation methodologies. To address these needs for robust AI safety evaluation, we introduce SAGE (Safety AI Generic Evaluation), an automated modular framework designed for customized and dynamic harm evaluations. SAGE employs prompted adversarial agents with diverse personalities based on the Big Five model, enabling system-aware multi-turn conversations that adapt to target applications and harm policies. We evaluate seven state-of-the-art LLMs across three applications and harm policies. Multi-turn experiments show that harm increases with conversation length, model behavior varies significantly when exposed to different user personalities and scenarios, and some models minimize harm via high refusal rates that reduce usefulness. We also demonstrate policy sensitivity within a harm category where tightening a child-focused sexual policy substantially increases measured defects across applications. These results motivate adaptive, policy-aware, and context-specific testing for safer real-world deployment.

大模型安全多轮评估对抗测试个性化风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。