arXiv:2510.05942cs.CLcs.AI2025-10中稿 · as a poster at *SE…被引 3

用思维链+大模型评分,评估20个大模型的道德对齐情况。

EvalMORAAL: Interpretable Chain-of-Thought and LLM-as-Judge Evaluation for Moral Alignment in Large Language Models

  • 构建透明思维链框架,结合概率评分与直接打分。
  • 西方模型对齐度达0.82,非西方仅0.61,差距显著。
  • 通过模型互评发现348处分歧,验证评估可靠性。

我们提出EvalMORAAL,一个透明的思维链(CoT)框架,通过日志概率和直接评分两种方法,结合模型作为评判者的同行评审,评估20个大语言模型在道德对齐方面的表现。评估基于世界价值观调查(WVS,55个国家,19个主题)和皮尤全球态度调查(PEW,39个国家,8个主题)。使用EvalMORAAL,顶尖模型与调查结果高度一致(WVS中皮尔逊相关系数$r \= 0.90$)。然而,区域差异明显:西方地区平均相关系数为0.82,非西方地区为0.61,绝对差距达0.21,表明存在持续的区域对齐鸿沟。框架包含三部分:(1)统一评分方法实现公平比较;(2)带自一致性检查的结构化思维链协议;(3)基于数据驱动阈值的模型互评机制,识别出348处冲突。同行评审一致性与WVS对齐度相关($r=0.74, p<.001$),PEW中为0.39(不显著),支持自动化质量控制的有效性。研究展示文化敏感型AI的进步,也揭示跨区域应用的挑战。

原文摘要 · Abstract (English)

We present EvalMORAAL, a transparent chain-of-thought (CoT) framework that uses two scoring methods (log-probabilities and direct ratings) plus a model-as-judge peer review to evaluate moral alignment in 20 large language models. We assess models on the World Values Survey (55 countries, 19 topics) and the PEW Global Attitudes Survey (39 countries, 8 topics). With EvalMORAAL, top models align closely with survey responses (Pearson's $r \approx 0.90$ on WVS). Yet we find a clear regional difference: Western regions average $r=0.82$ while non-Western regions average $r=0.61$ (a 0.21 absolute gap), indicating a persistent regional alignment gap. Our framework adds three parts: (1) two scoring methods for all models to enable fair comparison, (2) a structured CoT protocol with self-consistency checks, and (3) a model-as-judge peer review that flags 348 conflicts using a data-driven threshold. Peer agreement relates to WVS survey alignment ($r=0.74$, $p<.001$; PEW $r=0.39$, n.s.), supporting automated quality checks. These results show real progress toward culture-aware AI while highlighting open challenges for use across regions.

道德对齐思维链大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。