arXiv:2602.22697cs.CLcs.AI2026-02被引 1

让对话代理在用户体验与成本间找到平衡,提升真实场景下的服务效率。

Reinforcing Real-world Service Agents: Balancing Utility and Cost in Task-oriented Dialogue

  • 将对话任务建模为多粒度强化学习,通过用户驱动策略探索
  • 引入成本感知的多轮策略优化,在用户满意度与开销间实现帕累托最优
  • 适用于真实商业场景,尤其适合需控制成本的服务型智能体

大型语言模型的快速发展推动对话系统从聊天机器人向通用智能体演进。然而,如何在共情沟通与预算敏感决策之间有效权衡仍是开放挑战。现有方法难以捕捉这种复杂的策略权衡,为此我们提出InteractCS-RL框架,将任务导向型对话重构为多粒度强化学习过程。首先构建以用户为中心的交互框架,提供高保真训练环境,使智能体能动态探索多样化策略。其次提出成本感知的多轮策略优化(CMPO),采用混合优势估计策略,结合生成过程积分与PID-Lagrangian成本控制器,引导策略有效探索用户奖励与全局成本约束之间的帕累托边界。在定制化真实业务场景的大量实验表明,InteractCS-RL在三个评估维度上显著优于其他基线。进一步在工具-代理-用户交互基准上的评估验证了其在多领域中的鲁棒性。

原文摘要 · Abstract (English)

The rapid evolution of Large Language Models (LLMs) has accelerated the transition from conversational chatbots to general agents. However, effectively balancing empathetic communication with budget-aware decision-making remains an open challenge. Since existing methods fail to capture these complex strategic trade-offs, we propose InteractCS-RL, a framework that reframes task-oriented dialogue as a multi-granularity reinforcement learning process. Specifically, we first establish a User-centric Interaction Framework to provide a high-fidelity training gym, enabling agents to dynamically explore diverse strategies with persona-driven users. Then, we introduce Cost-aware Multi-turn Policy Optimization (CMPO) with a hybrid advantage estimation strategy. By integrating generative process credits and employing a PID-Lagrangian cost controller, CMPO effectively guides the policy to explore Pareto boundary between user reward and global cost constraints. Extensive experiments on customized real business scenarios demonstrate that InteractCS-RL significantly outperform other baselines across three evaluation dimensions. Further evaluation on tool-agent-user interaction benchmarks verify InteractCS-RL robustness across diverse domains.

对话系统强化学习成本优化智能代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。