构建多语言多模态电商评测基准,真实还原购物场景复杂性。
Towards Reliable Evaluation of Large Language Models for Multilingual and Multimodal E-Commerce Applications
- 基于真实用户查询与交易日志构建评测集,覆盖37项任务
- 引入专家审核的半自动流程,确保答案质量与可扩展性
- 支持七种语言,含五种东南亚低资源语言,适合跨语言研究
大型语言模型在通用NLP基准上表现优异,但在专业化领域能力仍待探索。现有电商评测如EcomInstruct、ChineseEcomQA、eCeLLM和Shopping MMLU存在任务多样性不足(缺乏产品推荐与售后问题)、模态单一(缺少多模态数据)、数据为合成或精选、且仅聚焦英语与中文等问题,导致从业者缺乏可靠工具评估模型在复杂真实购物场景中的表现。本文提出EcomEval,一个全面的多语言多模态电商评测基准。该基准涵盖6个类别、37项任务(含8项多模态任务),主要源自真实客户查询与交易日志,反映实际业务交互中的噪声与异构性。为保证参考答案的质量与可扩展性,采用半自动流程:由大模型生成候选回答,再经50多位具备电商与多语言专长的专家审校修改。通过不同规模模型的平均评分定义各题目与类别的难度等级,实现挑战导向与细粒度评估。EcomEval覆盖七种语言,包括五种低资源东南亚语言,填补了以往研究在多语言视角上的空白。
原文摘要 · Abstract (English)
Large Language Models (LLMs) excel on general-purpose NLP benchmarks, yet their capabilities in specialized domains remain underexplored. In e-commerce, existing evaluations-such as EcomInstruct, ChineseEcomQA, eCeLLM, and Shopping MMLU-suffer from limited task diversity (e.g., lacking product guidance and after-sales issues), limited task modalities (e.g., absence of multimodal data), synthetic or curated data, and a narrow focus on English and Chinese, leaving practitioners without reliable tools to assess models on complex, real-world shopping scenarios. We introduce EcomEval, a comprehensive multilingual and multimodal benchmark for evaluating LLMs in e-commerce. EcomEval covers six categories and 37 tasks (including 8 multimodal tasks), sourced primarily from authentic customer queries and transaction logs, reflecting the noisy and heterogeneous nature of real business interactions. To ensure both quality and scalability of reference answers, we adopt a semi-automatic pipeline in which large models draft candidate responses subsequently reviewed and modified by over 50 expert annotators with strong e-commerce and multilingual expertise. We define difficulty levels for each question and task category by averaging evaluation scores across models with different sizes and capabilities, enabling challenge-oriented and fine-grained assessment. EcomEval also spans seven languages-including five low-resource Southeast Asian languages-offering a multilingual perspective absent from prior work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。