arXiv:2602.07376cs.CL2026-02Conference of the …被引 2

测试大模型对不同人群安全观的敏感度,发现可实现兼顾多样性和评估可靠性。

Do Large Language Models Reflect Demographic Pluralism in Safety?

  • 在提示词层面建模14类安全领域,保留人口统计信息并扩充低资源领域。
  • 使用多模型零样本评估,达成高可靠性(ICC=0.87)与低人群敏感度(DS=0.12)。
  • 适合关注伦理多样性、模型公平性与安全评估方法的研究者。

大语言模型的安全性本应体现多元化的道德规范、文化期待与人口背景差异,但现有对齐数据集如 ANTHROPIC-HH 与 DICES 依赖人口结构单一的标注者,忽视了不同群体对安全认知的差异。为弥补这一缺口,Demo-SafetyBench 在提示词层面直接建模人口多样性,将价值框架与回答解耦。第一阶段:利用 Mistral 7B-Instruct-v0.3 将 DICES 提示重新分类至 14 个安全领域(源自 BEAVERTAILS),保留人口统计元数据,并通过 Llama-3.1-8B-Instruct 与 SimHash 去重技术扩充低资源领域,最终生成 43,050 条样本。第二阶段:采用 LLM-as-Raters 模式(Gemma-7B、GPT-4o、LLaMA-2-7B),在零样本推理下评估多元敏感性。设定平衡阈值(delta = 0.5,tau = 10)后,实现高可靠性(ICC = 0.87)与低人口敏感度(DS = 0.12),证明多元安全评估既可规模化又具人口鲁棒性。

原文摘要 · Abstract (English)

Large Language Model (LLM) safety is inherently pluralistic, reflecting variations in moral norms, cultural expectations, and demographic contexts. Yet, existing alignment datasets such as ANTHROPIC-HH and DICES rely on demographically narrow annotator pools, overlooking variation in safety perception across communities. Demo-SafetyBench addresses this gap by modeling demographic pluralism directly at the prompt level, decoupling value framing from responses. In Stage I, prompts from DICES are reclassified into 14 safety domains (adapted from BEAVERTAILS) using Mistral 7B-Instruct-v0.3, retaining demographic metadata and expanding low-resource domains via Llama-3.1-8B-Instruct with SimHash-based deduplication, yielding 43,050 samples. In Stage II, pluralistic sensitivity is evaluated using LLMs-as-Raters-Gemma-7B, GPT-4o, and LLaMA-2-7B-under zero-shot inference. Balanced thresholds (delta = 0.5, tau = 10) achieve high reliability (ICC = 0.87) and low demographic sensitivity (DS = 0.12), confirming that pluralistic safety evaluation can be both scalable and demographically robust.

安全评估大模型多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。