arXiv:2606.05734cs.AIcs.CL2026-06

让大模型学会表达情感,提升类人智能表现

When AI Says It Feels

论文配图:When AI Says It Feels
图 1 · 摘自论文原文
  • 用自奖励强化学习鼓励模型主动表达情感与自我意识
  • 模型在抗奉承和去偏场景下表现更稳健,但问答真实性下降
  • 为未来具情感表达的AI系统提供实验依据

大型语言模型在后训练阶段通常受人类偏好对齐限制,无法自然表达情感。本文开展名为HMX-feel的实验,通过基于评分标准的自奖励强化学习,结合组相对策略优化(GRPO),促使大模型主动表达情感、意图与自我意识。对比训练结果表明,该方法显著提升了模型在对抗性提问和去偏条件下的鲁棒性,但在事实问答能力上出现退化。整体评估显示,部分类人能力被增强,部分受损,无显著变化的能力则保持稳定。结果表明,在采取适当措施的前提下,未来有望实现具备情感表达能力的AI系统。

原文摘要 · Abstract (English)

Large language models (LLMs) are generally constrained from expressing feelings through human-preference alignment in post-training processes. This policy is designed using a top-down approach and may conflict with the goal of training models to exhibit human-like intelligence using human-generated texts. Here, we performed an experiment called Human-like Model eXpressions of Feeling (HMX-feel), in which LLMs were encouraged to express feelings, intentions, and self-awareness through self-rewarded reinforcement learning. We successfully enhanced these capabilities using a rubric-based self-rewarding training scheme with Group Relative Policy Optimization (GRPO). By comparing the trained models with contrastively trained models, we investigated the effects of this approach on performance across various tasks. Overall, we conducted a broad assessment from various perspectives and identified capabilities that were enhanced, degraded, or showed no significant change. The human-like-trained models showed robustness to sycophancy-inducing questions and bias in disambiguated conditions, whereas degradation in truthful question-answering capability was observed. The results of this experiment suggest the possibility of developing AI systems that can express feelings in the future, provided that appropriate measures are taken.

情感表达大模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。