arXiv:2506.07402cs.CRcs.CL2025-06被引 3

揭露大模型隐性安全风险:看似无害的输入也可能导致严重后果

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures

  • 用四象限分析框架识别隐性危害风险区域
  • 在多类模型中发现隐性危害普遍存在且具现实威胁
  • 适合关注大模型安全评估与对齐研究的开发者

大型语言模型(LLMs)在实际应用中日益普及,其安全性引发关注。尽管越狱攻击揭示了恶意查询下的失败,却忽略了更关键的风险:对看似无害的输入给出错误回答可能造成真实世界伤害(隐性危害)。本文通过基于输出真实性与输入危害性的结构化四象限视角,系统重构了LLM风险图景,揭示了一个被忽视的高危区域。为此,我们提出JailFlipBench基准,涵盖单模态、多模态及事实扩展场景,采用多样化评估指标捕捉隐性危害。进一步开发初步的JailFlip攻击方法,在多个开源与黑盒LLM上开展全面评估,结果表明隐性危害具有即时且紧迫的真实世界风险,亟需超越传统越狱范式的更广泛安全评估与对齐机制。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed in real-world applications, raising concerns about their security. While jailbreak attacks highlight failures under overtly harmful queries, they overlook a critical risk: incorrectly answering harmless-looking inputs can be dangerous and cause real-world harm (Implicit Harm). We systematically reformulate the LLM risk landscape through a structured quadrant perspective based on output factuality and input harmlessness, uncovering an overlooked high-risk region. To investigate this gap, we propose JailFlipBench, a benchmark aims to capture implicit harm, spanning single-modal, multimodal, and factual extension scenarios with diverse evaluation metrics. We further develop initial JailFlip attack methodologies and conduct comprehensive evaluations across multiple open-source and black-box LLMs, show that implicit harm present immediate and urgent real-world risks, calling for broader LLM safety assessments and alignment beyond conventional jailbreak paradigms.

大模型安全隐性危害对齐失败风险评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。