研究大模型如何被道德话术影响,揭示其伦理可塑性差异。
Moral Persuasion in Large Language Models: Evaluating Susceptibility and Ethical Alignment
- 用对话角色模拟道德说服,测试模型改变认知的能力。
- 不同大小同公司模型表现差异显著,说明伦理对齐非线性。
- 适合关注AI伦理安全与可控性的研究者阅读。
我们探究大型语言模型(LLMs)在提示下改变初始决策并与其对齐于既定伦理框架的可能。研究包含两项实验:第一项评估模型在道德模糊情境下的可说服性,通过说服者代理(Persuader Agent)尝试修改基础代理(Base Agent)的初始判断;第二项考察模型对预设伦理框架的响应能力,引导其采纳基于哲学理论的价值观。结果表明,大模型可在道德敏感场景中被有效说服,说服成功率受模型类型、情境复杂度及对话长度影响。值得注意的是,同公司但不同规模的模型表现出显著差异,凸显其伦理可塑性的不一致性。
原文摘要 · Abstract (English)
We explore how large language models (LLMs) can be influenced by prompting them to alter their initial decisions and align them with established ethical frameworks. Our study is based on two experiments designed to assess the susceptibility of LLMs to moral persuasion. In the first experiment, we examine the susceptibility to moral ambiguity by evaluating a Base Agent LLM on morally ambiguous scenarios and observing how a Persuader Agent attempts to modify the Base Agent's initial decisions. The second experiment evaluates the susceptibility of LLMs to align with predefined ethical frameworks by prompting them to adopt specific value alignments rooted in established philosophical theories. The results demonstrate that LLMs can indeed be persuaded in morally charged scenarios, with the success of persuasion depending on factors such as the model used, the complexity of the scenario, and the conversation length. Notably, LLMs of distinct sizes but from the same company produced markedly different outcomes, highlighting the variability in their susceptibility to ethical persuasion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。