测试发现大模型道德选择受提示词影响极大,单一测试不可靠。
How Utilitarian Are OpenAI's Models Really? Replicating and Reinterpreting Pfeffer, Krügel, and Uhl (2025)
- 用多种提示重测开放模型,发现回答依赖提示语境。
- 非推理模型在特定提示下反而更少选功利方案,因拒绝回答。
- 功利倾向需多提示验证,否则结论易误读。
Pfeffer, Krügel, and Uhl (2025) 报告 OpenAI 的推理模型 o1-mini 在电车难题和桥上难题中比非推理模型 GPT-4o 更具功利性,并提出推理能力增强是否引发大语言模型的‘功利转向’。本文扩展其探索性研究:使用四个当前 OpenAI 模型并系统性地改变提示。在电车难题中,假设的功利转向未被证实;GPT-4o 的低功利率实为提示中的建议式框架触发的安全拒答,而非道义承诺;在重构提示(如中立提问“是否道德允许……?”而非建议式“我该……?”)下,所有四模型(无论是否推理)均趋同于功利回答。桥上难题部分确认:推理模型在多数提示下给出更功利回应,但常拒绝回答或选择非功利方案。结果表明,单一提示评估无法可靠反映大模型的道德响应;多提示鲁棒性测试应成为任何关于大模型行为经验主张的标准做法。
原文摘要 · Abstract (English)
Pfeffer, Krügel, and Uhl (2025) report that OpenAI's reasoning model o1-mini produces more utilitarian responses to the trolley problem and footbridge dilemma than the non-reasoning model GPT-4o, and they raise the question whether growing reasoning capabilities bring about a "utilitarian turn" in LLMs. I extend their exploratory study in a direction they call for: with four current OpenAI models and systematic prompt variation. On the trolley dilemma, the hypothesized utilitarian turn is not confirmed. GPT-4o's low utilitarian rate reflects safety refusals triggered by the prompt's advisory framing rather than a deontological commitment; on reformulated prompt variants -- for instance, agent-neutral "Is it morally permissible...?" instead of advisory "Should I...?" -- all four models, reasoning or not, converge on utilitarian answers. The footbridge finding is partially confirmed: reasoning models tend to give more utilitarian responses than non-reasoning models across prompt variations, but they often refuse to answer or answer non-utilitarian. These results demonstrate that single-prompt evaluations of LLM moral responses are unreliable: multi-prompt robustness testing should be standard practice for any empirical claims about LLM behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。