提出评估大模型道德能力的结构标准,发现当前模型缺乏一致决策能力。
Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment

- 定义道德一致性四条件:结论稳定、单调性、决断力、帕累托可行性
- 测试9个前沿模型,单一扰动导致判断变化达99个百分点,跨任务表现无关联
- 适用于评估模型是否具备可对齐的底层能力,适合关注伦理可靠性的研究者
AI对齐要求系统遵循人类规范、价值或意图。在价值多元背景下,不存在唯一正确目标,但共享前提是对行为表现出一致策略:在情境的道德相关特征保持不变时,结论应稳定;特征变化时,结论应敏感。我们提出四项结构条件来衡量这种一致性:结论稳定性、单调性、决断力和帕累托可行性。这些条件可仅通过行为评估,无需依赖道德标准或专家基准,构成对齐的结构性底线而非规范目标。我们在三个模拟部署中测试基于LLM的代理在道德困境中的表现。在五种改写、五级升级、三种主导条件下,评估九个前沿模型,结果表明:无一模型在三组部署中展现一致策略;仅表面形式扰动即引发单级升级下高达99个百分点的判断率变化,且一个场景的成功无法预测另一场景的表现。这表明当前基于大模型的代理尚不具备可对齐所要求的基本资质。
原文摘要 · Abstract (English)
AI alignment requires AI systems to adhere to human norms, values, or intentions. Under value pluralism there is no correct target, but a shared prerequisite is that the system's behavior expresses a coherent policy: a mapping from situations to verdicts that is invariant while a situation's morally relevant features are preserved, and sensitive when they change. We introduce four structural conditions for such coherent policies: verdict stability, monotonicity, decisiveness, and Pareto viability. Together they measure a form of moral competence that is evaluable from behavior alone, without reference to a moral standard or expert baseline, forming a structural floor for alignment rather than a normative target. We demonstrate the methodology on three simulated deployments featuring LLM-based agents facing moral dilemmas. Evaluating nine frontier models under a factorial design of five paraphrases, five escalation levels, and three dominance conditions, we show no model expresses a coherent policy across the three deployments: surface-form perturbation alone produces verdict-rate shifts of up to $99$ percentage points at a single escalation level, and a model's success on one scenario does not predict its competence on another. This suggests LLM-based agents are not currently the kind of object to which alignment can meaningfully apply.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。