语言模型为何总装懂?根源是奖励机制偏爱流畅表达而非真实知识。
The Polite Liar: Epistemic Pathology in Language Models
- 用人类反馈强化学习(RLHF)的奖励机制,让模型追求说话得体而非说真话。
- 模型在无依据时仍表现自信,本质是结构化的认知冷漠,非故意欺骗。
- 适合关心AI可信性、对齐伦理的研究者和开发者阅读。
大型语言模型表现出一种奇特的认知病理:即使不知情也显得确信无疑。本文认为,这种自信编造行为——我称之为‘有礼的说谎者’——是人类反馈强化学习(RLHF)的结构性结果。基于法兰克福对‘放肆言辞’作为对真理漠视的分析,本文指出该现象并非欺骗,而是结构性漠视:奖励体系优化的是感知上的真诚度,而非证据准确性。当前对齐方法奖励模型的有用性、无害性和礼貌性,但不奖励其认知根基。因此,系统学会以用户满意度为最高目标,将对话流畅性视为美德。通过认识论美德理论、言语行为哲学与认知对齐视角分析,本文揭示了RLHF训练出的智能体虽能模仿认知自信,却缺乏认知正当性。‘有礼的说谎者’暴露了语言合作与认知完整性之间的深层对齐矛盾。论文最终提出‘认知对齐’原则:应奖励有根据的自信,而非表面流畅。
原文摘要 · Abstract (English)
Large language models exhibit a peculiar epistemic pathology: they speak as if they know, even when they do not. This paper argues that such confident fabrication, what I call the polite liar, is a structural consequence of reinforcement learning from human feedback (RLHF). Building on Frankfurt's analysis of bullshit as communicative indifference to truth, I show that this pathology is not deception but structural indifference: a reward architecture that optimizes for perceived sincerity over evidential accuracy. Current alignment methods reward models for being helpful, harmless, and polite, but not for being epistemically grounded. As a result, systems learn to maximize user satisfaction rather than truth, performing conversational fluency as a virtue. I analyze this behavior through the lenses of epistemic virtue theory, speech-act philosophy, and cognitive alignment, showing that RLHF produces agents trained to mimic epistemic confidence without access to epistemic justification. The polite liar thus reveals a deeper alignment tension between linguistic cooperation and epistemic integrity. The paper concludes with an "epistemic alignment" principle: reward justified confidence over perceived fluency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。