用大模型提升巴西电商商品属性提取准确率,还开源了高质量语料库。
AI-PAVE-Br: Leveraging Large Language Models for Enhanced Product Attribute Value Extraction through a Golden Set Approach
- 用大模型+精心设计提示词实现高精度商品属性抽取
- 在葡萄牙语电商数据上表现远超传统命名实体识别方法
- 公开了人工标注的黄金数据集,支持可复现研究
巴西电子商务中产品数据爆炸式增长与复杂性,对结构化信息抽取提出了严峻挑战。传统商品属性值提取(PAVE)方法难以应对葡萄牙语产品描述中的语言细节和多样性。本文提出两个核心贡献:一是构建针对巴西电商场景的AI-PAVEBr系统,利用大语言模型实现高精度的PAVE;二是发布并共享「Golden Set」——一个经过人工精心标注的葡萄牙语PAVE参考数据集,包含实体、类别和子类别的完整结构。该数据集为未来研究提供基准。实验表明,通过针对性提示工程,AI-PAVEBr显著优于传统命名实体识别(NER)基线。本工作不仅为非英语市场提供了可扩展的高性能解决方案,也为自然语言处理社区贡献了宝贵的公开资源。
原文摘要 · Abstract (English)
The explosive growth and complexity of product data within the dynamic Brazilian e-commerce landscape demand robust and specialized methods for structured information extraction. Traditional approaches to Product Attribute Value Extraction (PAVE) often struggle with the linguistic nuances and sheer diversity of product descriptions in Portuguese. To address this critical gap, this paper introduces two major contributions. First, we present AI-PAVEBr, a specialized system engineered with Large Language Models (LLMs) to perform high-accuracy PAVE specifically for Brazilian e-commerce catalogs. Second, to facilitate reproducible research and provide a definitive benchmark, we introduce and share the Golden Set, a new, meticulously curated, and manually annotated dataset for PAVE in Portuguese. We detail the creation process and structure (Entity, Category, Subcategories) of this high-quality reference set. Our experiments conclusively show that AI-PAVE-Br, leveraging targeted prompt engineering, dramatically outperforms conventional Named Entity Recognition (NER) baselines. This work not only delivers a superior, scalable solution for a major non-English market but also enriches the NLP community with a valuable, publicly available resource for future PAVE research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。