MUSE让中文对话模拟更真实连贯,支持多领域长期交互。
MUSE: Multi-Domain Chinese User Simulation via Self-Evolving Profiles and Rubric-Guided Alignment

- 通过自我进化用户画像逐步优化行为一致性
- 多轮强化学习结合评分模型,提升长对话表现
- 专为中文多领域设计,适合对话系统评估与训练
用户模拟器对交互式AI系统的可扩展训练与评估至关重要。然而,现有方法常依赖浅层用户画像,难以维持长期互动中的人物一致性,且多局限于英文或单领域场景。我们提出MUSE,一个面向多领域中文用户的仿真框架,旨在生成类人、可控且行为一致的回应。首先,提出迭代画像自进化(IPSE)机制,通过对比模拟轨迹与真实对话行为中的差异进行推理和优化。其次,采用角色反转监督微调提升局部响应的真实感与自然表达。为实现细粒度行为对齐,进一步训练基于评分标准的奖励模型,并融入评分引导的多轮强化学习,从对话层面优化模拟器,增强长时程行为一致性。实验表明,MUSE在话语级与会话级评估中均持续优于强基线,生成的回应更具真实性、连贯性与人物一致性。
原文摘要 · Abstract (English)
User simulators are essential for the scalable training and evaluation of interactive AI systems. However, existing approaches often rely on shallow user profiling, struggle to maintain persona consistency over long interactions, and are largely limited to English or single-domain settings. We present MUSE, a multi-domain Chinese user simulation framework designed to generate human-like, controllable, and behaviorally consistent responses. First, we propose Iterative Profile Self-Evolution (IPSE), which gradually optimizes user profiles by comparing and reasoning discrepancies between simulated trajectories and real dialogue behaviors. We then apply Role-Reversal Supervised Fine-Tuning to improve local response realism and human-like expression. To enable fine-grained behavioral alignment, we further train a specialized rubric-based reward model and incorporate it into rubric-guided multi-turn reinforcement learning, which optimizes the simulator at the dialogue level and enhances long-horizon behavioral consistency. Experiments show that MUSE consistently outperforms strong baselines in both utterance-level and session-level evaluations, generating responses that are more realistic, coherent, and persona-consistent over extended interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。