arXiv:2603.06065cs.IR2026-03

用强化学习训练更可靠的对话购物助手,兼顾准确性和说服力。

ChatShopBuddy: Towards Reliable Conversational Shopping Agents via Reinforcement Learning

  • 分层奖励建模,合理处理多种目标间的依赖关系。
  • 动态对比策略优化,提升响应质量与操作效率平衡。
  • 在真实场景中优于更大模型,稳定性更强。

对话式购物助手是大语言模型驱动代理的重要应用场景,但如何有效应用后训练强化学习(RL)进行优化仍缺乏研究。本文针对真实场景中的购物代理优化问题,探索了多目标协同的强化学习方法,需同时满足客观指标(商品正确性)、主观品质(说服力)、结果奖励(最终回复质量)和过程奖励(工具使用效率)。为此,我们构建了 SmartShopBench 基准,涵盖多样购物意图,并采用分层评估体系将复杂质量要求分解为可度量层级。在此基础上,提出分层奖励建模(HRM),通过条件门控结构反映不同奖励类型的逻辑依赖。为进一步提升训练效率,设计动态对比策略优化(DCPO),基于奖励与推理长度动态选择轨迹。大量实验表明,所提出的 RL 训练代理 ChatShopBuddy 在多个指标上持续优于依赖通用推理的大模型,表现出更强的稳定性而非仅峰值性能。本工作为强化学习在真实对话代理中的应用提供了重要实践指导。

原文摘要 · Abstract (English)

Conversational shopping agents represent a critical consumer-facing application of Large Language Model (LLM)-powered agents, yet how to effectively apply post-training Reinforcement Learning (RL) to optimize such agents remains underexplored. This work investigates RL-based optimization for shopping agents in real-world scenarios, where agents must simultaneously satisfy multiple interdependent objectives spanning objective metrics (product correctness), subjective qualities (persuasiveness), outcome rewards (final response quality), and process rewards (tool efficiency). We present a complete methodology to address this challenge. Specifically, we first construct SmartShopBench, a benchmark that captures diverse shopping intents with a hierarchical evaluation that decomposes complex quality requirements into measurable levels. Building on this evaluation framework, we design Hierarchical Reward Modeling (HRM) to structure mixed reward types through conditional gating that reflects their logical dependencies. To enable efficient training, we further propose Dynamic Contrastive Policy Optimization (DCPO), which balances response quality with operational efficiency through dynamic trajectory selection based on reward and reasoning length. Extensive experiments demonstrate that our RL-trained agent, namely ChatShopBuddy, consistently outperforms larger models relying on generic reasoning, achieving superior stability rather than merely higher peaks. Our work provides valuable guidance for applying RL to real-world conversational agents.

对话系统强化学习购物代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。