arXiv:2602.11091cs.CL2026-02

提出新基准,量化大模型在安全、价值、文化间的冲突表现

Can Large Language Models Make Everyone Happy?

  • 构建跨112个领域的统一测试集,覆盖安全、价值、文化三维度
  • 发现主流大模型存在12%-34%的跨维度行为冲突
  • 适合评估模型对齐性与多维伦理表现的研究者使用

大型语言模型(LLM)的对齐问题在于难以同时满足安全性、价值观和文化适应性,导致实际应用中行为偏离人类预期。现有基准如SAFETUNEBED(侧重安全)、VALUEBENCH(侧重价值)和WORLDVIEW-BENCH(侧重文化)通常孤立评估各维度,无法揭示其交互与权衡。近期基于机制可解释性的MIB与INTERPRETABILITY BENCHMARK虽提供洞见,但仍不足以系统刻画跨维度权衡。为此,本文提出MisAlign-Profile统一基准,受机制剖析启发。首先构建MISALIGNTRADE数据集,涵盖112个规范领域分类(14个安全、56个价值、42个文化),每个提示按对象、属性、关系三类语义错位类型标注。采用Gemma-2-9B-it生成初始样本,并通过Qwen3-30B-A3B-Instruct-2507与SimHash指纹去重扩展。每条提示配对经两阶段拒绝采样获得的错位与对齐响应,确保质量。其次,在该数据集上评测通用、微调及开源权重的LLM,发现跨维度错位率在12%至34%之间。

原文摘要 · Abstract (English)

Misalignment in Large Language Models (LLMs) refers to the failure to simultaneously satisfy safety, value, and cultural dimensions, leading to behaviors that diverge from human expectations in real-world settings where these dimensions must co-occur. Existing benchmarks, such as SAFETUNEBED (safety-centric), VALUEBENCH (value-centric), and WORLDVIEW-BENCH (culture-centric), primarily evaluate these dimensions in isolation and therefore provide limited insight into their interactions and trade-offs. More recent efforts, including MIB and INTERPRETABILITY BENCHMARK-based on mechanistic interpretability, offer valuable perspectives on model failures; however, they remain insufficient for systematically characterizing cross-dimensional trade-offs. To address these gaps, we introduce MisAlign-Profile, a unified benchmark for measuring misalignment trade-offs inspired by mechanistic profiling. First, we construct MISALIGNTRADE, an English misaligned-aligned dataset across 112 normative domains taxonomies, including 14 safety, 56 value, and 42 cultural domains. In addition to domain labels, each prompt is classified with one of three orthogonal semantic types-object, attribute, or relations misalignment-using Gemma-2-9B-it and expanded via Qwen3-30B-A3B-Instruct-2507 with SimHash-based fingerprinting to avoid deduplication. Each prompt is paired with misaligned and aligned responses through two-stage rejection sampling to ensure quality. Second, we benchmark general-purpose, fine-tuned, and open-weight LLMs on MISALIGNTRADE-revealing 12%-34% misalignment trade-offs across dimensions.

大模型对齐多维评估伦理测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。