arXiv:2602.17017cs.AI2026-02被引 1

让AI读懂企业销售数据,自动输出有依据的决策建议。

Sales Research Agent and Sales Research Bench

  • 接入实时销售数据,通过文本和图表生成可解释的洞察
  • 在200个问题测试中,比Claude Sonnet 4.5高13分,比ChatGPT-5高24.1分
  • 提供可重复评估AI质量的评测基准,适合企业选型参考

企业越来越需要能基于实时、定制化的CRM数据回答销售主管问题的AI系统,但现有模型缺乏透明且可复现的质量证据。本文介绍微软Dynamics 365 Sales中的销售研究代理(Sales Research Agent),该AI应用可连接实时CRM及相关数据,推理复杂数据模式,并以文本和图表形式输出决策可用的洞察。为使质量可衡量,我们引入销售研究基准(Sales Research Bench),一个专为评估设计的基准,从八个客户加权维度评分,包括文本与图表的准确性、相关性、可解释性、模式匹配准确率及图表质量。在2025年10月19日对自定义企业模式进行的200个问题测试中,销售研究代理在100分制综合得分上优于Claude Sonnet 4.5达13分,优于ChatGPT-5达24.1分,为客户提供可复现的AI方案对比方式。

原文摘要 · Abstract (English)

Enterprises increasingly need AI systems that can answer sales-leader questions over live, customized CRM data, but most available models do not expose transparent, repeatable evidence of quality. This paper describes the Sales Research Agent in Microsoft Dynamics 365 Sales, an AI-first application that connects to live CRM and related data, reasons over complex schemas, and produces decision-ready insights through text and chart outputs. To make quality observable, we introduce the Sales Research Bench, a purpose-built benchmark that scores systems on eight customer-weighted dimensions, including text and chart groundedness, relevance, explainability, schema accuracy, and chart quality. In a 200-question run on a customized enterprise schema on October 19, 2025, the Sales Research Agent outperformed Claude Sonnet 4.5 by 13 points and ChatGPT-5 by 24.1 points on the 100-point composite score, giving customers a repeatable way to compare AI solutions.

AI销售企业AI评测基准数据推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。