arXiv:2608.28610cs.AI2026-09

通过反馈机制测试大模型在道德决策中的行为变化。

TPvG: A Moral Decision Framework for Large Language Models from One-Shot to Sequential Feedback

论文配图:TPvG: A Moral Decision Framework for Large Language Models from One-Shot to Sequential Feedback
图 1 · 摘自论文原文
  • 设计五阶段任务,从单次决策到带反馈的连续选择
  • 模型对反馈响应差异大,且与人类模式不一致
  • 揭示大模型在交互场景下道德行为可能不稳定

现有大语言模型的道德评估通常只呈现孤立的道德情境并获取一次性决策,忽略了显著影响人类道德行为的重要因素:后果反馈。本文提出TPvG(Text-based Pain-versus-Gain),源自人类道德范式,将后果反馈融入日常道德困境——不伤害他人与最大化自我利益之间的权衡。TPvG包含五个道德决策任务,从低上下文的一次性选择逐步推进至带有明确接收者反馈的序列决策。实验结果表明,模型的道德决策强烈受决策格式(单次与序列)影响;显式接收者反馈对不同模型产生异质性影响。此外,模型对反馈的反应模式与人类参考模式存在显著偏离,暗示其决策过程可能与人类不同。这些发现凸显了评估大模型在高风险交互场景中道德行为稳定性的重要性。

原文摘要 · Abstract (English)

Existing LLM moral evaluations typically present models with isolated moral vignettes and elicit a single-shot decision, neglecting a factor known to profoundly influence human moral behavior: consequence feedback. We introduce TPvG (Text-based Pain-versus-Gain), adapted from a human moral paradigm, which embeds consequence feedback into an everyday moral dilemma of not harming others versus maximising self-gain. TPvG comprises five moral decision tasks, progressing from minimal-context one-shot choices to sequential decisions with explicit consequence feedback. Our results show that LLM moral decisions were strongly affected by decision format (one-shot versus sequential), and explicit receiver feedback produced heterogeneous effects across models. Furthermore, LLM responses to explicit receiver feedback diverged from the human reference pattern, suggesting potentially different decision processes. These findings highlight the need to evaluate whether LLM moral behavior remains stable in high-stakes interactive settings.

大模型道德推理反馈机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。