arXiv:2509.25137cs.AIcs.CL2025-09被引 15

让模型从真实对话中学习,实现更自然的个性化改进。

The Era of Real-World Human Interaction: RL from User Conversations

  • 通过用户后续对话自动修正输出,实现无标注反馈的强化学习。
  • 基于用户长期互动史建模人格,提升个性化与指令遵循能力。
  • 在真实对话数据上表现优于主流基线,适合追求真实交互体验的研究者。

我们认为,要实现模型持续优化和多维度对齐,未来模型必须从自然的人类交互中学习。当前对话模型依赖预标注、专家生成的人类反馈进行对齐。本文提出强化学习从人类交互(RLHI)新范式,直接从真实用户对话中学习。我们开发两种互补方法:(1) 基于用户引导重写的RLHI,利用用户自然语言的后续反馈修正模型不满意输出;(2) 基于用户奖励的RLHI,通过结合用户长期互动历史(即人格画像)的奖励模型进行学习。两者通过人格条件偏好优化,将长期用户人格与逐轮偏好关联。在WildChat数据集上训练的两种RLHI变体,在个性化和指令遵循上均优于强基线,且类似反馈也提升了推理基准表现。结果表明,自然的人类交互可为个性化对齐提供可扩展、有效的监督信号。

原文摘要 · Abstract (English)

We posit that to achieve continual model improvement and multifaceted alignment, future models must learn from natural human interaction. Current conversational models are aligned using pre-annotated, expert-generated human feedback. In this work, we introduce Reinforcement Learning from Human Interaction (RLHI), a paradigm that learns directly from in-the-wild user conversations. We develop two complementary methods: (1) RLHI with User-Guided Rewrites, which revises unsatisfactory model outputs based on users' natural-language follow-up responses, (2) RLHI with User-Based Rewards, which learns via a reward model conditioned on knowledge of the user's long-term interaction history (termed persona). Together, these methods link long-term user personas to turn-level preferences via persona-conditioned preference optimization. Trained on conversations derived from WildChat, both RLHI variants outperform strong baselines in personalization and instruction-following, and similar feedback enhances performance on reasoning benchmarks. These results suggest organic human interaction offers scalable, effective supervision for personalized alignment.

强化学习对话系统个性化真实数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。