解决销售对话中长期目标与即时语言质量的平衡难题
Harmonizing Dense and Sparse Signals in Multi-turn RL: Dual-Horizon Credit Assignment for Industrial Sales Agents
- 分时域信用分配,分离处理每轮与整段对话的奖励信号
- 转化率提升6.82%,重复率下降82.28%,身份识别率降27.35%
- 适合工业级销售对话系统优化,兼顾业绩与自然语言生成
工业级销售对话系统的优化需在长期商业目标(如转化率)与短期语言约束(如流畅性、合规性)之间取得平衡。传统强化学习常将异质目标合并为单一奖励,导致高量级会话级奖励压倒细微的轮次级信号,引发训练不稳定或奖励黑客问题。为此,本文提出双时域信用分配(DuCA)框架,通过核心机制——时域无关优势归一化(HIAN),在融合前分别对轮次级与会话级奖励的优势进行归一化,确保两类目标在策略更新中贡献均衡。基于高保真用户模拟器的大量实验表明,与最先进方法GRPO相比,DuCA实现转化率相对提升6.82%,句间重复率降低82.28%,身份检测率下降27.35%,显著提升了工业销售场景下战略表现与自然语言生成的协同能力。
原文摘要 · Abstract (English)
Optimizing large language models for industrial sales requires balancing long-term commercial objectives (e.g., conversion rate) with immediate linguistic constraints such as fluency and compliance. Conventional reinforcement learning often merges these heterogeneous goals into a single reward, causing high-magnitude session-level rewards to overwhelm subtler turn-level signals, which leads to unstable training or reward hacking. To address this issue, we propose Dual-Horizon Credit Assignment (DuCA), a framework that disentangles optimization across time scales. Its core, Horizon-Independent Advantage Normalization (HIAN), separately normalizes advantages from turn-level and session-level rewards before fusion, ensuring balanced gradient contributions from both immediate and long-term objectives to the policy update. Extensive experiments with a high-fidelity user simulator show DuCA outperforms the state-of-the-art GRPO baseline, achieving a 6.82% relative improvement in conversion rate, reducing inter-sentence repetition by 82.28%, and lowering identity detection rate by 27.35%, indicating a substantial improvement for an industrial sales scenario that effectively balances the dual demands of strategic performance and naturalistic language generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。