让大模型更会说服人:用多源奖励优化谈判对话效果
Teaching LLM to be Persuasive: Reward-Enhanced Policy Optimization for Alignment from Heterogeneous Rewards
- 融合偏好模型、人类评判和规则引擎的多源奖励机制
- 对话评分提升至4.63,优秀回复比例提高23.34个百分点
- 适合需要长程对话、合规性高的商业对话系统研发者
我们将大语言模型(LLM)部署为在线旅游平台(OTAs)中的业务拓展(BD)代理,用于具有说服力的价格谈判。该代理需遵循多阶段标准操作流程(SOP)并遵守严格约束(不夸大承诺、不生成幻觉),同时在长时间多轮对话中保持类人表现与高效性。我们提出奖励增强型策略优化(REPO),一种强化学习后训练方法,结合异构奖励:经偏好训练的奖励模型(RM)、以大模型为裁判(RJ)评估细微行为(如情感价值与SOP合规性)、以及基于规则的奖励函数(RF,主要为正则表达式)进行数值、格式与安全约束的确定性检查。在专家共识评估中(三名专家;30次真实对话与45个精心设计的错误案例),REPO将平均对话评分提升至4.63(较GRPO提升0.33),优秀回复占比达66.67%(提升23.34个百分点),坏例修复率达93.33%,其中75.56%为干净修复。在9,653条真实客户对话的生产环境A/B测试中(对比意图驱动对话系统),REPO使响应率提升12.14个百分点,任务成功率提升5.94个百分点(p<0.001)。
原文摘要 · Abstract (English)
We deploy large language models (LLMs) as business development (BD) agents for persuasive price negotiation in online travel agencies (OTAs). The agent must follow a multi-stage Standard Operating Procedure (SOP) and strict guardrails (no over-promising and no hallucinations), while remaining human-like and effective over long, multi-turn dialogues. We propose Reward-Enhanced Policy Optimization (REPO), a reinforcement learning post-training method that combines heterogeneous rewards: a preference-trained reward model (RM), an LLM-as-a-judge (RJ) for nuanced behaviors (e.g., emotional value and SOP compliance), and rule-based reward functions (RF) (mainly regex-based) for deterministic checks on numerics, formatting, and guardrails. In expert consensus evaluation (three human experts; 30 online conversations and 45 curated bad cases), REPO improves average dialogue rating to 4.63 (+0.33 over GRPO) and raises the share of conversations with at least one excellent response to 66.67% (+23.34 pp over GRPO), while achieving a 93.33% bad-case fix rate with 75.56% clean fixes. In a production A/B test on 9,653 real customer conversations (vs. an intent-driven dialogue system), REPO improves response rate by +12.14 pp and task success rate by +5.94 pp (p<0.001).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。