arXiv:2606.11232cs.CLcs.AI2026-06被引 1

测试大模型如何组合多个道德信号做出判断。

Every Act Has Its Price: Compressed Moral Composition in Frontier LLMs

论文配图:Every Act Has Its Price: Compressed Moral Composition in Frontier LLMs
图 1 · 摘自论文原文
  • 设计双阶段盲评竞技场,评估模型对复合道德情境的判断能力。
  • 模型对复合道德判断呈压缩关系,非简单加总,强度锚定明显。
  • 适合关注大模型伦理决策机制的研究者与开发者。

现有大模型道德评测多聚焦单一道德行为、价值或基础的偏好,虽有帮助但不完整。真实判断常需整合多个道德信号于同一选项中。本文提出「道德电车竞技场」(Moral Trolley Arena),一种两阶段盲式ELO评测框架,用于衡量大模型如何组合道德证据。单场景竞技场首先基于229个场景语料库,在五种道德基础理论(Moral Foundations Theory)基础上校准各道德行为;复合竞技场则在受控强度网格下,将校准后的行为组合成双行为道德项,并测量其复合偏好。在十款前沿模型中,复合判断主要由分量行为强度预测,但关系呈现系统性压缩而非简单叠加。模型还表现出非加性强度锚定、成分控制后仍存在的基础特异性残差,以及跨厂商高度收敛的复合偏好表面。结果表明,道德审计应关注道德证据的组合规则,而不仅限于孤立行为的排序。

原文摘要 · Abstract (English)

Existing LLM moral benchmarks usually ask which isolated moral act, value, or foundation a model prefers. This is useful but incomplete. Realistic judgments often require a model to combine several moral signals within the same option. We introduce **Moral Trolley Arena**, a two-stage blind ELO benchmark for measuring how LLMs compose moral evidence. The single-scene arena first calibrates individual moral acts from a 229-scenario corpus across five Moral Foundations Theory foundations; the composite arena then combines calibrated acts into two-act moral items over a controlled intensity grid and measures the resulting composite preferences. Across ten frontier models, composite judgments are largely predicted by component act strength, but the relation is consistently compressed rather than simply additive. Models also show non-additive intensity anchoring, bounded foundation-specific residuals after component control, and highly convergent composite preference surfaces across providers. These results suggest that moral audits should measure composition rules for moral evidence, not only rankings over isolated acts.

道德推理大模型评估组合判断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。