首个评估大模型真实销售能力的多轮对话基准,测进度也看成交。
Sell More, Play Less: Benchmarking LLM Realistic Selling Skill

- 构建含3万+脚本的真实销售场景,支持中英文双语测试。
- 用评分模型与意图分类器自动评估销售进展和购买意愿,相关性达0.86。
- 发现中文最强模型仅相当于初级销售员,跨语言表现差异大。
销售对话需在信息不对称下进行多轮目标导向说服,对大模型构成挑战。现有对话评测很少衡量交易推进与结果。本文提出SalesLLM基准,基于金融与消费品领域的真实应用,包含30,074个脚本配置和1,805个精心设计的多轮情景,支持难度与角色可调。我们构建全自动评估流程:(i) 用LLM裁判评估销售进程,(ii) 用微调BERT分类器判断对话结束时的购买意图。为提升模拟真实性,基于8,000+真人参与的销售对话训练用户模型CustomerLM,将角色错位率从GPT-4o的17.44%降至8.8%。该基准与人工评分高度一致(平均皮尔逊相关系数r=0.86;标注者间一致性Krippendorff's alpha=0.86,每模型500次标注)。在15个主流大模型中,中文最强模型表现接近初级至中级人类销售人员,但未达专家水平,弱模型则低于此基准,跨语言一致性仍差。SalesLLM为面向结果的销售代理提供可扩展评测标准。代码与数据已开源。
原文摘要 · Abstract (English)
Sales dialogues require multi-turn, goal-directed persuasion under asymmetric incentives, which makes them a challenging setting for large language models (LLMs). Yet existing dialogue benchmarks rarely measure deal progression and outcomes. We introduce SalesLLM benchmark, a bilingual (ZH/EN) benchmark derived from realistic applications covering Financial Services and Consumer Goods, built from 30,074 scripted configurations and 1,805 curated multi-turn scenarios with controllable difficulty and personas. We propose a fully automatic evaluation pipeline that combines (i) an LLM judge for sales-process progress, and (ii) fine-tuned BERT classifiers for end-of-dialogue buying intent. To improve simulation fidelity, we train a user model, CustomerLM, with SFT and DPO on 8,000+ crowdworker-involved sales conversations, reducing role inversion from 17.44% (GPT-4o) to 8.8%. SalesLLM benchmark scores correlate strongly with human ratings (average Pearson r=0.86; inter-annotator agreement Krippendorff's alpha=0.86, 500 annotations per model). Across 15 mainstream LLMs, in Chinese the strongest models are competitive with typical (junior-to-intermediate) human salespeople -- not with sales experts -- while weaker ones fall below this baseline, and cross-lingual consistency remains poor. SalesLLM benchmark serves as a scalable benchmark for outcome-oriented sales agents. Our code and data are released at https://github.com/Bairong-Xdynamics/Benchmarking-LLM-Realistic-Selling-Skill
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。