arXiv:2505.09576cs.CYcs.AI2025-05中稿 · version被引 1

剖析强化学习中的人类反馈机制如何暗中塑造语言与人际互动的伦理问题。

Ethics and Persuasion in Reinforcement Learning from Human Feedback: A Procedural Rhetorical Approach

  • 用程序修辞理论分析RLHF背后的语言操控逻辑。
  • 揭示其可能固化语言霸权、加剧偏见、削弱学习语境。
  • 适合教育者、AI研究者及生成式AI使用者阅读。

自2022年以来,以ChatGPT和Claude为代表的生成式AI聊天机器人采用一种称为人类反馈强化学习(Reinforcement Learning from Human Feedback, RLHF)的技术,通过人工标注者的反馈对大语言模型(LLMs)进行微调。这一机制显著提升了模型输出质量,使交互内容更趋近于“人类表达”。然而,人机文本日益趋同,带来透明度、信任、偏见及人际关系等深层次的伦理、社会技术与教育影响。本文从程序修辞(procedural rhetoric)视角切入,分析当前由RLHF驱动的生成式聊天机器人所重构的核心过程:语言规范维护、信息获取方式以及社交关系期待。不同于以往聚焦生成内容说服力的研究,本文将批判焦点转向嵌入在模型训练机制中的深层说服逻辑。该理论探讨为人工智能伦理开辟新方向,指出算法流程可能强化主流语言使用、延续偏见、剥离学习语境,并侵蚀人际边界。因此,该研究对教育工作者、研究人员、学者及日益广泛的生成式AI用户具有重要参考价值。

原文摘要 · Abstract (English)

Since 2022, versions of generative AI chatbots such as ChatGPT and Claude have been trained using a specialized technique called Reinforcement Learning from Human Feedback (RLHF) to fine-tune language model output using feedback from human annotators. As a result, the integration of RLHF has greatly enhanced the outputs of these large language models (LLMs) and made the interactions and responses appear more "human-like" than those of previous versions using only supervised learning. The increasing convergence of human and machine-written text has potentially severe ethical, sociotechnical, and pedagogical implications relating to transparency, trust, bias, and interpersonal relations. To highlight these implications, this paper presents a rhetorical analysis of some of the central procedures and processes currently being reshaped by RLHF-enhanced generative AI chatbots: upholding language conventions, information seeking practices, and expectations for social relationships. Rhetorical investigations of generative AI and LLMs have, to this point, focused largely on the persuasiveness of the content generated. Using Ian Bogost's concept of procedural rhetoric, this paper shifts the site of rhetorical investigation from content analysis to the underlying mechanisms of persuasion built into RLHF-enhanced LLMs. In doing so, this theoretical investigation opens a new direction for further inquiry in AI ethics that considers how procedures rerouted through AI-driven technologies might reinforce hegemonic language use, perpetuate biases, decontextualize learning, and encroach upon human relationships. It will therefore be of interest to educators, researchers, scholars, and the growing number of users of generative AI chatbots.

AI伦理语言模型强化学习程序修辞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。