arXiv:2601.08951cs.CYcs.AI2026-01被引 5

构建多元人类判断的AI危害评估基准,揭示分歧根源。

PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm

  • 设计双维度基准:危害程度与共识度,捕捉真实人类分歧
  • 150个高分歧提示,1.5万份标注,涵盖心理与人口特征
  • 发现即时风险与具体伤害提升感知危害,个体差异影响判断

当前AI安全框架常将危害性视为二元判断,难以处理人类存在实质性分歧的边缘案例。为构建更包容的系统,需超越共识,理解分歧发生的位置与原因。本文提出PluriHarms基准,系统研究人类对AI危害判断的两个核心维度——危害轴(良性到有害)与共识轴(一致到分歧)。该可扩展框架生成能反映多样AI危害与人类价值观的提示,重点聚焦高分歧场景,并通过人类数据验证。基准包含150个提示,15,000份来自100名标注者的人类评分,附带人口统计、心理特质及提示层面的危害行为、影响与价值特征。分析表明,涉及即时风险和具体伤害的提示会加剧危害感知;标注者特质(如毒性经历、教育背景)及其与提示内容的交互作用可解释系统性分歧。我们在PluriHarms上评估了多个AI安全模型与对齐方法,发现个性化显著提升对人类判断的预测能力,但仍有较大改进空间。本工作通过显式关注价值多样性与分歧,为突破‘一刀切’安全范式、迈向多元共治的AI安全提供了原则性基准。

原文摘要 · Abstract (English)

Current AI safety frameworks, which often treat harmfulness as binary, lack the flexibility to handle borderline cases where humans meaningfully disagree. To build more pluralistic systems, it is essential to move beyond consensus and instead understand where and why disagreements arise. We introduce PluriHarms, a benchmark designed to systematically study human harm judgments across two key dimensions -- the harm axis (benign to harmful) and the agreement axis (agreement to disagreement). Our scalable framework generates prompts that capture diverse AI harms and human values while targeting cases with high disagreement rates, validated by human data. The benchmark includes 150 prompts with 15,000 ratings from 100 human annotators, enriched with demographic and psychological traits and prompt-level features of harmful actions, effects, and values. Our analyses show that prompts that relate to imminent risks and tangible harms amplify perceived harmfulness, while annotator traits (e.g., toxicity experience, education) and their interactions with prompt content explain systematic disagreement. We benchmark AI safety models and alignment methods on PluriHarms, finding that while personalization significantly improves prediction of human harm judgments, considerable room remains for future progress. By explicitly targeting value diversity and disagreement, our work provides a principled benchmark for moving beyond "one-size-fits-all" safety toward pluralistically safe AI.

AI安全人类判断价值对齐评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。