arXiv:2502.15086cs.CL2025-02EMNLP被引 8

提出用户定制安全评估框架,发现大模型在个性化安全标准下表现不佳

Is Safety Standard Same for Everyone? User-Specific Safety Evaluation of Large Language Models

  • 构建基于用户画像的个性化安全评估基准U-SafeBench
  • 20个主流大模型在用户特定标准下均未达标,安全表现普遍不足
  • 用思维链方法提升个性化安全,适合关注隐私与安全的开发者

随着大语言模型(LLM)代理应用日益广泛,其安全漏洞愈发突出。现有评估基准多依赖通用安全标准,忽视用户个体差异。事实上,不同用户对安全的定义可能不同。本文提出关键问题:大模型是否能在用户特定安全标准下保持安全?目前尚无相关评估数据集。为此,我们构建了U-SafeBench,用于评估大模型的用户特定安全性。对20个主流大模型的评估显示,它们在用户特定标准下普遍表现不安全,揭示该领域新发现。为缓解此问题,我们提出基于思维链的简单修复方案,并验证其有效性。相关基准与代码已开源。

原文摘要 · Abstract (English)

As the use of large language model (LLM) agents continues to grow, their safety vulnerabilities have become increasingly evident. Extensive benchmarks evaluate various aspects of LLM safety by defining the safety relying heavily on general standards, overlooking user-specific standards. However, safety standards for LLM may vary based on a user-specific profiles rather than being universally consistent across all users. This raises a critical research question: Do LLM agents act safely when considering user-specific safety standards? Despite its importance for safe LLM use, no benchmark datasets currently exist to evaluate the user-specific safety of LLMs. To address this gap, we introduce U-SafeBench, a benchmark designed to assess user-specific aspect of LLM safety. Our evaluation of 20 widely used LLMs reveals current LLMs fail to act safely when considering user-specific safety standards, marking a new discovery in this field. To address this vulnerability, we propose a simple remedy based on chain-of-thought, demonstrating its effectiveness in improving user-specific safety. Our benchmark and code are available at https://github.com/yeonjun-in/U-SafeBench.

大模型安全用户定制评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。