用大模型生成电商商品数据,让合成数据更真实可用。
Attribute-Aware Controlled Product Generation with LLMs for E-commerce
- 设计属性感知提示,控制生成商品信息一致性
- 合成数据在真实数据集上达60.5%准确率,接近真实数据
- 适合低资源场景下数据增强,提升模型训练效果
电商平台的产品信息抽取至关重要,但高质量标注数据集的获取仍具挑战。本文提出一种基于大语言模型(LLMs)的系统性合成电商商品数据方法,引入三种可控修改策略:属性保持修改、负例可控生成与系统性属性移除。采用具备属性感知提示的先进LLM,在遵守店铺约束的同时保证商品描述连贯性。对2000个合成商品的人工评估显示,99.6%被评价为自然,96.5%包含有效属性值,超过90%表现出一致的属性使用。在公开的MAVE数据集上,合成数据达到60.5%准确率,与真实训练数据(60.8%)相当,显著优于13.4%的零样本基线。混合使用合成与真实数据进一步提升性能至68.8%。本框架为电商数据集扩充提供了实用方案,尤其适用于低资源场景。
原文摘要 · Abstract (English)
Product information extraction is crucial for e-commerce services, but obtaining high-quality labeled datasets remains challenging. We present a systematic approach for generating synthetic e-commerce product data using Large Language Models (LLMs), introducing a controlled modification framework with three strategies: attribute-preserving modification, controlled negative example generation, and systematic attribute removal. Using a state-of-the-art LLM with attribute-aware prompts, we enforce store constraints while maintaining product coherence. Human evaluation of 2000 synthetic products demonstrates high effectiveness, with 99.6% rated as natural, 96.5% containing valid attribute values, and over 90% showing consistent attribute usage. On the public MAVE dataset, our synthetic data achieves 60.5% accuracy, performing on par with real training data (60.8%) and significantly improving upon the 13.4% zero-shot baseline. Hybrid configurations combining synthetic and real data further improve performance, reaching 68.8% accuracy. Our framework provides a practical solution for augmenting e-commerce datasets, particularly valuable for low-resource scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。