评估大模型卖家代理在电商砍价中的多轮意图理解能力
Evaluating Multi-Turn Bargain Skills in LLM-Based Seller Agent
- 基于心智理论构建逐轮评估框架,标注买家意图
- 涵盖622个品类、3014个任务的大规模基准数据集
- 适合研究对话智能与电商自动化谈判的开发者
在二手交易市场中,多轮讨价还价是买卖双方互动的核心环节。大语言模型可作为卖家代理,在既定商业约束下与买家进行谈判。关键能力在于持续追踪并准确理解谈判过程中累积的买家意图,直接影响议价效果。本文提出一种用于评估电商对话中卖家代理议价能力的多轮评估框架,测试其提取和跟踪买家意图的能力。主要贡献包括:(1) 构建覆盖622个品类、9,892件商品、3,014个任务的大规模电商议价基准;(2) 基于心智理论(Theory of Mind)设计逐轮评估框架,包含标注的买家意图,突破仅依赖结果指标的局限;(3) 开发自动化管道,从海量对话数据中可靠提取意图。
原文摘要 · Abstract (English)
In online second-hand marketplaces, multi-turn bargaining is a crucial part of seller-buyer interactions. Large Language Models (LLMs) can act as seller agents, negotiating with buyers on behalf of sellers under given business constraints. A critical ability for such agents is to track and accurately interpret cumulative buyer intents across long negotiations, which directly impacts bargaining effectiveness. We introduce a multi-turn evaluation framework for measuring the bargaining ability of seller agents in e-commerce dialogues. The framework tests whether an agent can extract and track buyer intents. Our contributions are: (1) a large-scale e-commerce bargaining benchmark spanning 622 categories, 9,892 products, and 3,014 tasks; (2) a turn-level evaluation framework grounded in Theory of Mind (ToM) with annotated buyer intents, moving beyond outcome-only metrics; and (3) an automated pipeline that extracts reliable intent from massive dialogue data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。