arXiv:2607.20471cs.AI2026-07

评测大模型在真实销售场景中的个性化说服能力,发现现有模型效果已达瓶颈。

Benchmarking the Personalization Capabilities of Large Language Models

论文配图:Benchmarking the Personalization Capabilities of Large Language Models
图 1 · 摘自论文原文
  • 基于贝叶斯劝说框架构建可复现的个性化评估范式
  • 在6279条企业成功案例上测试,顶尖模型无法区分成功与失败沟通
  • 实测中48%生成内容被销售人员视为立即可用,专家评分相关性达0.82

个性化指在发送方、渠道和时间固定的情况下,调整信息以促使特定接收者采取行动,是心理学与营销中的经典二元问题。大语言模型通过生成受接收者状态影响的消息变体,突破了传统检索-排序方法的库存限制,但其在经典意义上的个性化效果尚不明确。现有评测仅关注发送方自身适应性,未考察生成内容是否真正影响第三方。本文将Kamenica和Gentzkow(2011)的贝叶斯劝说框架适配至生成式代理,并在销售场景中实现,该场景可记录每条外联与其对应行动。我们发布SDR-Bench,一个包含6279条跨22个行业、约200家企业的客户成功故事的公开语料库,通过时间约束模拟避免未来数据泄露。在前沿大模型与深度研究代理中,均观察到个性化表现趋于平稳,对《财富》百强科技企业样本,无模型能统计区分成功与失败外联。一项实地部署测试显示,48%的模型生成内容被12名专业销售代表评为立即可用,资深专家评分相关性为Pearson 0.82。我们公开发布SDR-Arena与SDR-Bench,以支持大规模可复现的生成式个性化研究。

原文摘要 · Abstract (English)

Personalization, the act of varying a message to induce action from a specific receiver while keeping sender, channel, and time fixed, has a long tradition in psychology and marketing as a two-party problem in which sender and receiver have independent objectives. Large language models remove the bounded-inventory constraint of classical retrieval-and-ranking approaches by generating a continuum of message variants conditioned on inferred receiver state, raising the question of how well current models perform personalization in the classical sense. Existing LLM personalization benchmarks measure sender-side adaptation, in which the receiver is the same user the model is serving. The two-party question, whether a generated message induces its intended action in a third party, has been investigated only through A/B tests and small-scale human studies that cannot be re-run against a new model on demand. We adapt the Bayesian Persuasion framework of Kamenica and Gentzkow (2011) to generative agents and instantiate the formulation in sales, where receiver actions are routinely logged against the outreach that induced them. We release SDR-Bench, a public corpus of 6,279 customer success stories spanning 22 industries and approximately 200 enterprises, served through a temporally constrained simulation that prevents future-data leakage. Across frontier LLMs and deep-research agents, we observe a consistent personalization plateau and on a Fortune 100 tech cohort no model statistically separates successful from unsuccessful outreach. A field deployment with 12 professional sales representatives validates the framework, with 48 percent of model-generated content rated immediately useful and senior-expert agreement at Pearson 0.82. We release SDR-Arena and SDR-Bench publicly to support reproducible study of generative personalization at scale.

大模型评测个性化销售生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。