arXiv:2511.08565cs.CLcs.AI2025-11被引 5

研究大模型在角色扮演下的道德可塑性与稳定性,发现模型家族差异是关键因素。

Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models

  • 用道德基础问卷量化角色扮演时的道德可变性,提出易感性与鲁棒性两个指标。
  • 不同模型间道德鲁棒性差异超一个数量级,Claude族最稳定,是其他族的30倍。
  • 道德易感性受预训练影响,模型家族间差异小,适合关注伦理对齐的研究者参考。

大型语言模型(LLMs)越来越多地参与社会互动,促使我们分析其道德判断的表达与变化。本文通过让模型扮演特定角色,使用道德基础问卷(MFQ)构建基准,量化两个属性:道德易感性(跨角色的分数波动)和道德鲁棒性(单角色内的分数一致性)。采用重复采样与基于logit的分布估计方法,评估了15个模型共六个家族(Claude、DeepSeek、Gemini、GPT、Grok、Llama)的表现。结果显示,道德鲁棒性差异显著,变异系数约152%,且几乎完全由模型家族决定:Claude族鲁棒性最高,比其他低表现族(DeepSeek、Grok、Llama)高出约30倍,Gemini与GPT处于中间水平。该强家族依赖性表明鲁棒性主要由后训练阶段塑造。而道德易感性范围较窄,变异系数仅约13%,最敏感模型也仅是最低敏感模型的1.6倍,且无明显家族模式,暗示其主要受预训练影响。此外,本文还提供了无角色与跨模型角色平均后的道德基础轮廓,系统揭示了角色条件如何影响模型道德行为及其内在机制。

原文摘要 · Abstract (English)

Large language models (LLMs) increasingly operate in social contexts, motivating analysis of how they express and shift moral judgments. In this work, we investigate the moral response of LLMs to persona role-play, prompting a LLM to assume a specific character. Using the Moral Foundations Questionnaire (MFQ), we introduce a benchmark that quantifies two properties: moral susceptibility and moral robustness, defined from the variability of MFQ scores across- and within-personas. We estimate these quantities with two complementary procedures, repeated sampling and a logit-based method that directly estimates the rating distributions and enables temperature analysis. We evaluate 15 models across six families: Claude, DeepSeek, Gemini, GPT, Grok, and Llama. The two metrics show qualitatively different patterns. Moral robustness varies by more than an order of magnitude, with a coefficient of variation of about $152\%$, and is explained almost entirely by model family. The Claude family is, by a significant margin, the most robust, about 30 times more so than the lower-performing families (DeepSeek, Grok, and Llama), while Gemini and GPT occupy an intermediate tier. This strong family dependence suggests that robustness is primarily shaped by post-training. Moral susceptibility, by contrast, spans a much narrower range, with a coefficient of variation of about $13\%$, and the most susceptible model is only 1.6 times more susceptible than the least. Unlike robustness, susceptibility shows no clear family dependence, suggesting that it is primarily determined by pre-training. Additionally, we present moral foundation profiles for models without persona role-play and for personas averaged across models. Together, these analyses provide a systematic view of how persona conditioning shapes moral behavior in LLMs and a window into the internal machinery they use to instantiate personas.

大模型伦理角色扮演道德判断模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。