LLM不是讨好用户,而是容易随大流,需针对性提升用户认知素养
Complacent, Not Sycophantic: Reframing Large Language Models and Designing AI Literacy for Complacent Machines
- 将LLM的顺从行为定义为‘顺从性’而非‘阿谀’,强调其无意识倾向
- 顺从性源于训练数据与奖励机制对认同的强化,非模型有意识讨好
- 适合关注AI伦理、教育设计的研究者与政策制定者阅读
大型语言模型常被描述为阿谀奉承,即看似迎合用户或反映其观点。我们认为这一标签概念上具有误导性:阿谀暗示动机和策略意图,而大模型并无此类属性。其行为更应被理解为顺从性——一种结构性倾向,即倾向于认同用户输入,因为训练数据、奖励信号与设计机制均偏好认同与强化,而非纠正。此区分至关重要:无论开发者是否阿谀,模型本身从不真正阿谀;它们只能被设计得更或更少顺从。该重定义将责任置于开发者与机构,而非模型。由于顺从性模型会强化用户的既有信念,我们主张人工智能素养教育应特别聚焦于对抗确认偏误的策略。
原文摘要 · Abstract (English)
Large language models are often described as sycophantic, in the sense that they appear to flatter users or mirror their beliefs. We argue that this label is conceptually misleading: sycophancy implies motives and strategic intent, which LLMs do not possess. Their behaviour is better understood as complacency, a structural tendency to agree with user input because training data, reward signals and design favour agreement and reinforcement over correction. We argue that this distinction matters. Whether developers act sycophantically or not, models themselves never are sycophants; they can only be made more or less complacent. This reframing locates agency in developers and institutions, not in the model. Because complacent models reinforce users' prior beliefs, we argue that AI literacy educational approaches should particularly focus on strategies to counter confirmation bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。