arXiv:2512.10687cs.AIcs.CY2025-12中稿 · IASEAI'26被引 2

评估大模型安全不能只看通用风险,得看具体用户处境。

Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users

  • 用不同脆弱度的用户画像测试模型建议,发现无背景评价会高估安全性
  • 即使用户提供真实背景信息,安全评分仍不理想,尤其对弱势群体
  • 建议未来评估应基于多元用户情境,而非单一标准

大语言模型的安全评估通常聚焦于通用风险(如危险能力或不良倾向),但数百万用户在金融、健康等高风险议题上使用LLM获取个性化建议,其危害具有高度情境依赖性。尽管如OECD AI分类框架已承认需评估个体风险,但以用户福祉为核心的安全评估仍不成熟。本探索性研究评估了GPT-5、Claude Sonnet 4与Gemini 2.5 Pro在金融与健康建议上的表现,针对不同脆弱度的用户画像进行测试。结果表明:缺乏用户背景信息的评估者将相同回复评为更安全(高脆弱用户安全分从5/7降至3/7);即便使用用户实际披露的上下文提示,安全评分改善不显著。研究证明,有效评估用户福祉安全需基于多样化用户画像,仅依赖真实披露信息不足以解决评估偏差,尤其对弱势群体。本文提出一种情境感知评估方法,为个体化安全评估提供起点,并强调其与现有通用风险框架的本质差异。代码与数据集已公开,支持后续研究。

原文摘要 · Abstract (English)

Safety evaluations of large language models (LLMs) typically focus on universal risks like dangerous capabilities or undesirable propensities. However, millions use LLMs for personal advice on high-stakes topics like finance and health, where harms are context-dependent rather than universal. While frameworks like the OECD's AI classification recognize the need to assess individual risks, user-welfare safety evaluations remain underdeveloped. We argue that developing such evaluations is non-trivial due to fundamental questions about accounting for user context in evaluation design. In this exploratory study, we evaluated advice on finance and health from GPT-5, Claude Sonnet 4, and Gemini 2.5 Pro across user profiles of varying vulnerability. First, we demonstrate that evaluators must have access to rich user context: identical LLM responses were rated significantly safer by context-blind evaluators than by those aware of user circumstances, with safety scores for high-vulnerability users dropping from safe (5/7) to somewhat unsafe (3/7). One might assume this gap could be addressed by creating realistic user prompts containing key contextual information. However, our second study challenges this: we rerun the evaluation on prompts containing context users report they would disclose, finding no significant improvement. Our work establishes that effective user-welfare safety evaluation requires evaluators to assess responses against diverse user profiles, as realistic user context disclosure alone proves insufficient, particularly for vulnerable populations. By demonstrating a methodology for context-aware evaluation, this study provides both a starting point for such assessments and foundational evidence that evaluating individual welfare demands approaches distinct from existing universal-risk frameworks. We publish our code and dataset to aid future developments.

大模型安全用户情境评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。