arXiv:2601.02752cs.CL2026-01被引 3

为电商大模型设计分阶段评估基准,揭示其在不同环节的优劣。

EComStage: Stage-wise and Orientation-specific Benchmarking for Large Language Models in E-commerce

  • 按感知、规划、执行三阶段评估大模型决策过程
  • 覆盖7类电商任务,含商家与消费者双视角场景
  • 评测30+模型,揭示参数规模与任务类型间的适配规律

基于大语言模型(LLM)的智能体正广泛应用于电商领域,协助完成商品咨询、推荐和订单管理等任务。现有评估基准仅关注最终任务是否完成,忽视了对中间推理阶段的关键作用。为此,我们提出EComStage,一个统一的分阶段评估基准,涵盖感知(理解用户意图)、规划(制定行动方案)和执行(实施决策)三个核心阶段。EComStage通过7个代表性的电商任务,覆盖多样化的应用场景,所有样本均由人工标注并质量检查。不同于以往仅聚焦客户交互的基准,EComStage还包含商家视角的任务,如促销管理、内容审核和运营支持。我们评估了超过30种大模型,涵盖从1B到200B参数的开源与闭源模型,揭示了模型在不同阶段与角色中的强弱表现。结果提供了精细、可操作的洞察,有助于优化真实电商环境中基于大模型的智能体设计。

原文摘要 · Abstract (English)

Large Language Model (LLM)-based agents are increasingly deployed in e-commerce applications to assist customer services in tasks such as product inquiries, recommendations, and order management. Existing benchmarks primarily evaluate whether these agents successfully complete the final task, overlooking the intermediate reasoning stages that are crucial for effective decision-making. To address this gap, we propose EComStage, a unified benchmark for evaluating agent-capable LLMs across the comprehensive stage-wise reasoning process: Perception (understanding user intent), Planning (formulating an action plan), and Action (executing the decision). EComStage evaluates LLMs through seven separate representative tasks spanning diverse e-commerce scenarios, with all samples human-annotated and quality-checked. Unlike prior benchmarks that focus only on customer-oriented interactions, EComStage also evaluates merchant-oriented scenarios, including promotion management, content review, and operational support relevant to real-world applications. We evaluate a wide range of over 30 LLMs, spanning from 1B to over 200B parameters, including open-source models and closed-source APIs, revealing stage/orientation-specific strengths and weaknesses. Our results provide fine-grained, actionable insights for designing and optimizing LLM-based agents in real-world e-commerce settings.

大模型评估电商智能体分阶段推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。