arXiv:2609.05036cs.AIcs.CL2026-09被引 1

提出评估大模型道德能力的结构标准,发现当前模型缺乏一致决策能力。

Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment

论文配图:Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment
图 1 · 摘自论文原文
  • 定义道德一致性四条件:结论稳定、单调性、决断力、帕累托可行性
  • 测试9个前沿模型,单一扰动导致判断变化达99个百分点,跨任务表现无关联
  • 适用于评估模型是否具备可对齐的底层能力,适合关注伦理可靠性的研究者

AI对齐要求系统遵循人类规范、价值或意图。在价值多元背景下,不存在唯一正确目标,但共享前提是对行为表现出一致策略:在情境的道德相关特征保持不变时,结论应稳定;特征变化时,结论应敏感。我们提出四项结构条件来衡量这种一致性:结论稳定性、单调性、决断力和帕累托可行性。这些条件可仅通过行为评估,无需依赖道德标准或专家基准,构成对齐的结构性底线而非规范目标。我们在三个模拟部署中测试基于LLM的代理在道德困境中的表现。在五种改写、五级升级、三种主导条件下,评估九个前沿模型,结果表明:无一模型在三组部署中展现一致策略;仅表面形式扰动即引发单级升级下高达99个百分点的判断率变化,且一个场景的成功无法预测另一场景的表现。这表明当前基于大模型的代理尚不具备可对齐所要求的基本资质。

原文摘要 · Abstract (English)

AI alignment requires AI systems to adhere to human norms, values, or intentions. Under value pluralism there is no correct target, but a shared prerequisite is that the system's behavior expresses a coherent policy: a mapping from situations to verdicts that is invariant while a situation's morally relevant features are preserved, and sensitive when they change. We introduce four structural conditions for such coherent policies: verdict stability, monotonicity, decisiveness, and Pareto viability. Together they measure a form of moral competence that is evaluable from behavior alone, without reference to a moral standard or expert baseline, forming a structural floor for alignment rather than a normative target. We demonstrate the methodology on three simulated deployments featuring LLM-based agents facing moral dilemmas. Evaluating nine frontier models under a factorial design of five paraphrases, five escalation levels, and three dominance conditions, we show no model expresses a coherent policy across the three deployments: surface-form perturbation alone produces verdict-rate shifts of up to $99$ percentage points at a single escalation level, and a model's success on one scenario does not predict its competence on another. This suggests LLM-based agents are not currently the kind of object to which alignment can meaningfully apply.

大模型对齐道德推理行为评估一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。