用排序方法测大模型价值结构对齐,发现表面得分好也可能内部错位。
Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-Sorts

- 让人类和模型对140条道德陈述做九列强制分布排序,比对结构一致性。
- 12个模型在两温度下测试,显示跨家族差异和生成随机性敏感性。
- 揭示了全局高分下仍存在局部价值错位,适合评估伦理模型结构可靠性。
大语言模型在需要复杂道德推理和价值权衡的场景中日益普及。然而,现有评估多依赖个体行为指标,无法捕捉模型作为整体如何结构性地优先处理冲突价值。为此,我们提出一种基于Q方法的对称人-模型评估框架,测量价值结构对齐程度。在该协议下,人类与模型对同一组140条道德陈述进行九列强制分配排序;对模型,我们获取严格排序并确定性映射至排序桶。使用人类参考样本(N=35),我们建立了一个稳定且特定于该工具与样本的三因子参考几何。通过240次重复排序,在四个模型族中评估12个大模型,采用Procrustes相似性(ϕ)与基于RSA的Spearman相关系数(ρ)量化结构对齐。结果表明,不同模型族间存在显著异质性,模型对生成随机性敏感,且存在局部错位现象,说明良好全局分数可能掩盖深层区域偏差。尽管基于排名与桶的分析高度一致,但提示词措辞引入显著差异。最终,评估价值结构对齐可为传统逐项道德基准提供关键的结构性补充。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly deployed in contexts requiring complex moral reasoning and value trade-offs. However, existing evaluations typically rely on item-level behavioral metrics, which fail to capture how models structurally prioritize competing values as a cohesive system. To address this, we propose a symmetric human-LLM evaluation framework, grounded in Q methodology, to measure value-structure alignment. Under our protocol, humans and models sort an identical 140-item moral statement set into a shared nine-column forced distribution; for LLMs, we elicit strict rankings and deterministically map them to Q-sort buckets. Using a human reference sample ($N=35$), we establish a stable three-factor reference geometry specific to this instrument and sample. We evaluate 12 LLMs across four model families via 240 replicated Q-sorts at two temperature settings, quantifying structural alignment via Procrustes similarity ($ϕ$) and RSA-based Spearman correlation ($ρ$). Our results reveal significant cross-family heterogeneity, model-specific sensitivity to generation stochasticity and localized misalignment, which demonstrate that favorable global scores can obscure underlying regional distortions. While rank- and bucket-based analyses remain highly consistent, prompt phrasing introduces notable variance. Ultimately, assessing value-structure alignment provides a crucial structural complement to traditional itemwise moral benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。