arXiv:2507.05639cs.CL2025-07EMNLP被引 29

首个评估大模型在电商客服中多模态能力的基准测试

ECom-Bench: Can LLM Agent Resolve Real-World E-commerce Customer Support Issues?

  • 基于真实用户对话生成动态角色模拟,还原真实客服场景
  • 即使GPT-4o在复杂任务上通过率也仅10%-20%
  • 适合研究大模型在真实电商服务中的应用与改进

本文提出ECom-Bench,首个评估具备多模态能力的大语言模型代理在电商客服领域的基准框架。该框架基于真实电商交互中收集的人物画像信息,实现动态用户模拟,并构建源自真实电商对话的现实任务数据集。这些任务覆盖广泛商业场景,反映真实世界复杂性,具有高度挑战性。例如,即使是先进的GPT-4o模型,在本基准测试中的通过率(pass^3)也仅为10%-20%,凸显复杂电商场景带来的巨大难度。代码与数据已公开于https://github.com/XiaoduoAILab/ECom-Bench,以促进该领域进一步研究与发展。

原文摘要 · Abstract (English)

In this paper, we introduce ECom-Bench, the first benchmark framework for evaluating LLM agent with multimodal capabilities in the e-commerce customer support domain. ECom-Bench features dynamic user simulation based on persona information collected from real e-commerce customer interactions and a realistic task dataset derived from authentic e-commerce dialogues. These tasks, covering a wide range of business scenarios, are designed to reflect real-world complexities, making ECom-Bench highly challenging. For instance, even advanced models like GPT-4o achieve only a 10-20% pass^3 metric in our benchmark, highlighting the substantial difficulties posed by complex e-commerce scenarios. The code and data have been made publicly available at https://github.com/XiaoduoAILab/ECom-Bench to facilitate further research and development in this domain.

大模型代理电商客服多模态评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。