arXiv:2503.15990cs.CL2025-03被引 8

构建电商知识图谱评测基准,精准评估大模型事实性。

ECKGBench: Benchmarking Large Language Models in E-commerce Leveraging Knowledge Graph

  • 基于大规模知识图谱自动生成可靠问答数据
  • 通过最少输入输出token实现高效评估
  • 融合电商专家经验,揭示大模型知识边界

大型语言模型(LLMs)在多个自然语言处理任务中展现出强大能力,其在电商领域的潜力亦不容忽视,已应用于平台搜索、个性化推荐和客户服务等场景。然而,模型幻觉(hallucination)问题尤为突出,直接影响用户体验与收益。现有评估方法存在可靠性不足、成本高、缺乏领域专长等问题。为此,本文提出ECKGBench,一个专为评估电商领域语言模型能力设计的基准数据集。我们采用标准化流程,基于大规模知识图谱自动生成问题,确保评估可靠性;采用简单问答范式,以最少的输入输出令牌显著提升评估效率;并在人类标注、提示工程、负样本采样和验证环节融入丰富电商专业知识。此外,从新视角探索了大模型在电商领域的知识边界。通过对多个先进大模型在ECKGBench上的全面评估,提供了深入分析与实践洞察。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated their capabilities across various NLP tasks. Their potential in e-commerce is also substantial, evidenced by practical implementations such as platform search, personalized recommendations, and customer service. One primary concern associated with LLMs is their factuality (e.g., hallucination), which is urgent in e-commerce due to its significant impact on user experience and revenue. Despite some methods proposed to evaluate LLMs' factuality, issues such as lack of reliability, high consumption, and lack of domain expertise leave a gap between effective assessment in e-commerce. To bridge the evaluation gap, we propose ECKGBench, a dataset specifically designed to evaluate the capacities of LLMs in e-commerce knowledge. Specifically, we adopt a standardized workflow to automatically generate questions based on a large-scale knowledge graph, guaranteeing sufficient reliability. We employ the simple question-answering paradigm, substantially improving the evaluation efficiency by the least input and output tokens. Furthermore, we inject abundant e-commerce expertise in each evaluation stage, including human annotation, prompt design, negative sampling, and verification. Besides, we explore the LLMs' knowledge boundaries in e-commerce from a novel perspective. Through comprehensive evaluations of several advanced LLMs on ECKGBench, we provide meticulous analysis and insights into leveraging LLMs for e-commerce.

大模型评测电商知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。