arXiv:2511.00222cs.CLcs.AI2025-11NeurIPS被引 42

用强化学习让AI角色对话更稳定可信。

Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning

  • 设计三类自动评估指标,捕捉角色行为漂移
  • 多轮强化学习使角色不一致率降低55%以上
  • 适合需要稳定角色模拟的教育与心理场景

大型语言模型(LLMs)在治疗、教育和社交角色扮演等交互场景中被广泛用于模拟人类用户。然而,现成的LLMs常出现角色偏离、前后矛盾或脱离角色行为的问题。本文提出一个统一框架,用于评估和提升对话中的角色一致性。定义了三种自动度量:提示到语句一致性、语句到语句一致性、问答一致性,分别捕捉不同类型的角色漂移,并通过人工标注验证其有效性。将这些度量作为奖励信号,采用多轮强化学习对LLM进行微调,针对患者、学生和社交聊天伙伴三种角色。实验表明,该方法使角色不一致性降低超过55%,显著提升了模拟用户的连贯性与真实性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used to simulate human users in interactive settings such as therapy, education, and social role-play. While these simulations enable scalable training and evaluation of AI agents, off-the-shelf LLMs often drift from their assigned personas, contradict earlier statements, or abandon role-appropriate behavior. We introduce a unified framework for evaluating and improving persona consistency in LLM-generated dialogue. We define three automatic metrics: prompt-to-line consistency, line-to-line consistency, and Q&A consistency, that capture different types of persona drift and validate each against human annotations. Using these metrics as reward signals, we apply multi-turn reinforcement learning to fine-tune LLMs for three user roles: a patient, a student, and a social chat partner. Our method reduces inconsistency by over 55%, resulting in more coherent and faithful simulated users.

角色一致性强化学习对话系统LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。