用强化学习闭环优化销售线索排序,提升真实转化率。
SalesLoop: Reinforcement Learning from Performance Feedback for Sales Lead Ranking

- 基于实际转化结果构建位置与速度加权奖励函数。
- 列表级优化使排序准确率提升7.9%,关键指标提升15.8%。
- 适合关注落地效果的工业界模型迭代团队。
客户关系管理(CRM)系统中的线索排序面临长期挑战:离线精度高的模型在生产环境表现不佳。我们识别出三大根本差距:离线-在线指标不匹配、点对点与列表级目标错位、时间分布漂移。为此提出SalesLoop,一种建立模型预测与真实业务结果闭环反馈的强化学习框架。方法引入(1)性能感知奖励,融合转化结果、排序位置与转化速度;(2)判别式GRPO,将分组相对策略优化适配至判别式排序模型。在某新能源车企160天生产A/B测试中,覆盖1650万线索与280名销售专家,两个省级市场验证累计提升4.7%(p=0.047)和8.7%(p=0.002)。生产环境中,排名主干模型实现前10%召回率44.1%,高意向线索转化率达专员基线的2.3倍。
原文摘要 · Abstract (English)
Lead ranking in Customer Relationship Management (CRM) systems faces a persistent challenge: models achieving high offline accuracy often underperform in production. We identify three fundamental gaps responsible for this disconnect: offline-online metric mismatch, pointwise-listwise objective misalignment, and temporal distribution drift. To address these gaps, we propose SalesLoop, a reinforcement learning framework that establishes a closed feedback loop between model predictions and real-world business outcomes. Our approach introduces (1) a performance-aware reward that encodes conversion outcomes weighted by ranking position and conversion velocity, and (2) Discriminative GRPO, a listwise optimization objective that adapts Group Relative Policy Optimization to discriminative ranking models. SalesLoop improves NDCG@K by +7.9\% and P@K by +15.8\% over the strongest static baseline. A 160-day production A/B test at a New Energy Vehicle manufacturer, spanning 16.5M leads and 280 sales specialists across two provincial markets, validates statistically significant cumulative lift of +4.7\% ($p=0.047$) and +8.7\% ($p=0.002$). In production, the ranking backbone achieves Top-10\% recall of 44.1\% and surfaces high-intent leads at $2.3\times$ the conversion rate of specialist baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。