用多模态大模型自动评估电商搜索,省时省钱还高效
Retrieve, Annotate, Evaluate, Repeat: Leveraging Multimodal LLMs for Large-Scale Product Retrieval Evaluation
- 用多模态大模型生成每条查询的定制化标注指南
- 评估结果与人工标注相当,效率提升显著
- 适合需要大规模搜索系统质量控制的电商平台
在大规模电子商务场景中,对商品搜索系统进行规模化评估是一项关键但困难的任务,主要受限于高质量人工标注者数量有限。大型语言模型(LLMs)具备解决这一扩展性问题的潜力,可作为人类标注的可行替代方案,承担大部分标注任务。本文提出一种框架,利用多模态大模型实现:(i) 为每个查询生成定制化的标注指南;(ii) 执行后续的标注任务。该方法在某大型电商平台部署验证,结果表明其评估质量与人工标注相当,显著降低时间和成本,加快问题发现速度,为生产级检索系统的规模化质量控制提供了有效方案。
原文摘要 · Abstract (English)
Evaluating production-level retrieval systems at scale is a crucial yet challenging task due to the limited availability of a large pool of well-trained human annotators. Large Language Models (LLMs) have the potential to address this scaling issue and offer a viable alternative to humans for the bulk of annotation tasks. In this paper, we propose a framework for assessing the product search engines in a large-scale e-commerce setting, leveraging Multimodal LLMs for (i) generating tailored annotation guidelines for individual queries, and (ii) conducting the subsequent annotation task. Our method, validated through deployment on a large e-commerce platform, demonstrates comparable quality to human annotations, significantly reduces time and cost, facilitates rapid problem discovery, and provides an effective solution for production-level quality control at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。