arXiv:2512.00344cs.AI2025-12被引 4

让大模型学会根据用户性格实时调整对话风格,提升情感化交互质量。

Echo-N1: Affective RL Frontier

  • 基于用户实时性格推断,动态优化对话行为以匹配个性化偏好。
  • 在非可验证场景下实现显著且稳定的对话质量提升,超越基线模型与闭源产品。
  • 首次构建动态情绪智能评估体系,量化主观对话表现的改进。

大语言模型领域过去一年专注于强化学习在数学、编程和确定性推理等机器已擅长任务上的优化,却完全忽略了真正定义人类智能的领域:带有主观情感、人格敏感性的对话。这一领域常被视为难以形式化,不适合传统强化学习流程。本文证明其不仅可行,更是一个可解决且具有变革性的强化学习问题。我们提出首个框架,能在对话中实时推断用户人格,并据此优化模型行为以满足个性化偏好。不同于普遍认为强化学习在不可验证环境中会失效的观点,我们的方法在人机交互质量上实现了稳定、显著且戏剧性的提升。同时,我们引入首个动态情绪智能评估套件,用于量化这些进步。所提出的Echo-N1模型表现远超其基础版本,且优于闭源的Doubao 1.5 Character。这项工作确立了强化学习的新前沿:优化模型以应对对话中深层的主观性与人性维度。

原文摘要 · Abstract (English)

The LLM field has spent a year perfecting RL for tasks machines already excel at, math, code, and deterministic reasoning, while completely sidestepping the domain that actually defines human intelligence: subjective, emotionally grounded, personality sensitive conversation. This space has often been regarded as inherently subjective and challenging to formalize, making it appear unsuitable for conventional RL pipelines. We show that it is not only possible and it is a solvable and transformative RL problem. We propose the first framework that infers user personality on the fly and optimizes model behavior toward personalized conversational preferences. Contrary to the widespread belief that RL collapses in non-verifiable settings, our method produces consistent, robust, and dramatic improvements in humanlike interaction quality. We also introduce the first dynamic emotional intelligence evaluation suite to quantify these gains. Our model, which is introduced as Echo-N1, behaves far above its base version and outperforming the proprietary Doubao 1.5 Character. This work establishes a new frontier for RL: optimizing models for the deeply subjective, deeply human dimensions of conversation.

强化学习对话系统情感智能个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。