用人机协作迭代生成电商属性体系,提升搜索精准度。
BEATS: Bootstrapping E-commerce Attribute Taxonomies for Search through Iterative Human-AI Collaboration
- 通过多轮人机协作,从零构建产品属性分类体系。
- 在2,694个子类中生成67,277个属性,覆盖超540万商品。
- 显著改善搜索的筛选、排序与语义理解能力,适合本地化电商优化。
新兴市场电商平台常缺乏结构化产品属性,仅具备类别分类,导致搜索功能受限——无法实现细粒度筛选、查询理解下降、检索表征弱化。本文提出BEATS,一种基于大模型的人机协作框架,可从零构建产品属性分类体系。该框架包含两阶段生产流程:(1)模型开发者主动检查输出质量以过滤错误结果;(2)领域专家本地人员进行人工验证。系统采用迭代机制,根据每轮质检与标注反馈持续优化提示词,逐步提升属性质量。属性体系建立后,利用大模型对单个商品进行结构化属性标注,增强其上下文表征。丰富后的商品数据直接赋能搜索系统:支持细粒度属性筛选、为排序模型提供结构化特征、提升密集检索的语义表达。我们通过在属性增强数据上训练稠密检索模型,验证了其相较原始目录数据的持续性能提升。系统已部署于Rakuten Taiwan,涵盖9大类、2,694个子类,生成67,277个属性,超540万商品已完成属性标注,并计划全面覆盖全品类商品目录。
原文摘要 · Abstract (English)
E-commerce platforms in emerging markets often operate with underdeveloped product catalogs that contain only category taxonomies but lack structured attribute schemas. This absence of fine-grained product attributes limits search capabilities -- preventing faceted filtering, degrading query understanding, and weakening semantic representations used by search systems. We present BEATS, a human-in-the-loop LLM framework for bootstrapping product attribute taxonomies entirely from scratch. Our approach extends a multi-stage LLM generation pipeline with two critical production stages: (1) proactive quality checking by model developers to filter erroneous outputs, and (2) human annotation by domain-expert local staff to validate generated attributes. The framework operates iteratively -- prompts at each generation stage are refined based on quality check observations and annotator feedback across successive rounds, progressively improving attribute quality. Once the attribute taxonomy is established, we employ LLMs to perform structured attribute tagging on individual product items, enriching their contextual representations. The enriched catalog directly benefits multiple components of the search system: enabling granular attribute-based filtering, providing structured features for ranking models, and improving semantic representations for dense retrieval. We validate the generated taxonomy by training dense retrieval models on attribute-enriched product data, demonstrating consistent improvements over baselines using original catalog information. Our system has been deployed at Rakuten Taiwan, enriching 9 major categories spanning 2,694 sub-categories with 67,277 generated attributes, and over 5.4 million products have been tagged with the generated attributes, with plans to enrich the entire product catalog.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。