arXiv:2502.20196cs.CL2025-02KDD被引 12

构建中文电商概念评估基准,检验大模型真实业务理解能力。

ChineseEcomQA: A Scalable E-commerce Concept Evaluation Benchmark for Large Language Models

论文配图:ChineseEcomQA: A Scalable E-commerce Concept Evaluation Benchmark for Large Language Models
图 1 · 摘自论文原文
  • 聚焦基础电商概念,覆盖多样化任务场景。
  • 结合大模型与人工验证,确保评估结果可靠。
  • 适合研究电商大模型能力或落地应用的团队参考。

随着大语言模型在电商等领域的广泛应用,评估其领域能力的专用评测基准至关重要。现有基准面临两大挑战:任务类型多样且异构,难以统一评估;难以区分通用能力与电商特异性能力。为此,我们提出 extbf{ChineseEcomQA},一个面向基础电商概念的可扩展问答评测基准。该基准具备三大核心特性:聚焦基础概念、兼顾电商通用性与专业性。基础概念设计为适用于多种电商任务,有效应对任务多样性问题。通过平衡通用性与专属性,能精准验证模型在电商领域的理解深度。基准构建采用可扩展流程,融合大模型验证、检索增强生成(RAG)验证与严格人工标注。基于该基准,我们对主流大模型进行了全面评估,揭示了其在电商场景中的表现差异与潜在缺陷。期望 ChineseEcomQA 能推动未来领域专用评测的发展,促进大模型在电商中的更广泛落地。

原文摘要 · Abstract (English)

With the increasing use of Large Language Models (LLMs) in fields such as e-commerce, domain-specific concept evaluation benchmarks are crucial for assessing their domain capabilities. Existing LLMs may generate factually incorrect information within the complex e-commerce applications. Therefore, it is necessary to build an e-commerce concept benchmark. Existing benchmarks encounter two primary challenges: (1) handle the heterogeneous and diverse nature of tasks, (2) distinguish between generality and specificity within the e-commerce field. To address these problems, we propose \textbf{ChineseEcomQA}, a scalable question-answering benchmark focused on fundamental e-commerce concepts. ChineseEcomQA is built on three core characteristics: \textbf{Focus on Fundamental Concept}, \textbf{E-commerce Generality} and \textbf{E-commerce Expertise}. Fundamental concepts are designed to be applicable across a diverse array of e-commerce tasks, thus addressing the challenge of heterogeneity and diversity. Additionally, by carefully balancing generality and specificity, ChineseEcomQA effectively differentiates between broad e-commerce concepts, allowing for precise validation of domain capabilities. We achieve this through a scalable benchmark construction process that combines LLM validation, Retrieval-Augmented Generation (RAG) validation, and rigorous manual annotation. Based on ChineseEcomQA, we conduct extensive evaluations on mainstream LLMs and provide some valuable insights. We hope that ChineseEcomQA could guide future domain-specific evaluations, and facilitate broader LLM adoption in e-commerce applications.

电商大模型评测中文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。