arXiv:2608.09420cs.CL2026-08

让用户模拟器能精准控制对话意图,提升交互真实性。

Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation

论文配图:Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation
图 1 · 摘自论文原文
  • 将对话意图显式作为每轮指令,分离意图与语言表达
  • 在LMSYS-USP上实现86.6%意图准确率,领先24.3个百分点
  • 适合需要精确控制对话方向的智能助手训练与评估

用户模拟器广泛用于训练和评估交互式助手。生成下一轮用户回应本质上是多对一:相同用户画像和对话上下文可能有多种合理延续,对应不同局部交互意图。流畅的回应可能因意图错误(如接受而非修复)而误导对话。本研究的核心洞察是:可控用户模拟应将意图选择与语言表达方式分离。提出UserIDA(User Intent-Directive Alignment),将交互意图作为每轮显式指令。UserIDA定义六类意图接口,通过监督微调学习指令条件生成,并在群体强化学习中使用意图校准策略优化。奖励函数兼顾综合回应质量,确保违背意图的候选在混合组中排名低于合规选项。在LMSYS-USP数据集上,UserIDA达到86.6%意图准确率,比最强基线高出24.3个百分点,同时提升语义与风格相似性。在上下文干预测试中,91.7%的对话状态可实现至少四种目标意图,远超外部基线的22.9%。结果表明,每轮意图控制是用户模拟中响应保真度的有力补充。

原文摘要 · Abstract (English)

User simulators are widely used as scalable environments for training and evaluating interactive assistants. Generating the next user turn is inherently one-to-many: the same profile and dialogue context may support multiple plausible continuations with different local interaction intents. A fluent response may therefore advance the dialogue through an inappropriate intent, such as acceptance rather than repair. Our key insight is that controllable user simulation should separate which local interaction intent the next user turn should realize from how that intent is expressed in language. We introduce UserIDA (User Intent-Directive Alignment), which exposes interaction intent as an explicit per-turn directive. UserIDA defines a six-way intent interface, learns directive-conditioned generation through supervised fine-tuning, and uses intent-calibrated policy optimization during group-based reinforcement learning. The reward preserves composite response quality while ensuring that intent-violating candidates rank below compliant alternatives in mixed groups. On LMSYS-USP, UserIDA achieves 86.6\% intent accuracy, outperforming the strongest dedicated user-simulator baseline by 24.3 percentage points while improving semantic and stylistic similarity. In within-context interventions, it realizes at least four of the six target intents in 91.7\% of evaluated dialogue states, compared with 22.9\% for the strongest external baseline. These results establish per-turn intent control as a complementary dimension to response fidelity in user simulation.

用户模拟意图控制对话系统强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。