arXiv:2601.15330cs.CLcs.AI2026-01中稿 · ICASSP 2026被引 2

让大模型学会在模糊指令下主动求澄清,避免对话跑偏。

ICPO: Illocution-Calibrated Policy Optimization for Multi-Turn Conversation

  • 通过识别用户意图,训练模型在不确定时主动提问
  • 多轮对话准确率平均提升75%,单轮任务不退化
  • 适合需要高可靠性对话的智能客服、助手场景

多轮对话中,大语言模型常因早期错误假设陷入‘迷失’困境,尤其当用户初始指令模糊时。我们发现,标准后训练方法如可验证奖励强化学习(RLVR)会加剧此问题,因其奖励自信直接回答,导致模型过度自信并回避澄清。为此,我们提出意旨校准策略优化(ICPO),在训练数据中引入不明确提示,并将奖励信号与用户的言外之意关联,鼓励模型在面对歧义时表达不确定或主动询问。实验表明,ICPO显著提升模型谦逊性,多轮对话性能平均提升75%,同时保持单轮基准任务的稳定表现。该工作为构建更稳健、协作性强的对话式AI提供了可行路径。

原文摘要 · Abstract (English)

Large Language Models (LLMs) in multi-turn conversations often suffer from a ``lost-in-conversation'' phenomenon, where they struggle to recover from early incorrect assumptions, particularly when users provide ambiguous initial instructions. We find that standard post-training techniques like Reinforcement Learning with Verifiable Rewards (RLVR) exacerbate this issue by rewarding confident, direct answers, thereby inducing overconfidence and discouraging the model from seeking clarification. To address this, we propose Illocution-Calibrated Policy Optimization (ICPO), a novel training framework that sensitizes the model to instruction ambiguity. ICPO augments the training corpus with underspecified prompts and conditions the reward signal on the user's illocutionary intent, rewarding the model for expressing uncertainty or asking for clarification when faced with ambiguity. Experiments demonstrate that ICPO fosters appropriate humility, yielding a substantial average improvement of 75\% in multi-turn conversation, while preserving robust performance on single-turn benchmarks. Our work presents a practical path toward more robust and collaborative conversational AI that can better navigate the nuances of human interaction.

对话系统强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。