无需标注数据,用自动优化提示词评估电商商品质量
Auto prompting without training labels: An LLM cascade for product quality assessment in e-commerce catalogs
- 通过级联生成与优化提示词,实现无监督商品质量评估
- 相比传统方法,精确率和召回率提升8%-10%
- 将专家工作时间从5.1小时降至3分钟,适合多语言场景
我们提出一种无需训练的级联自动提示框架,用于在电商目录中评估商品质量。该系统不依赖训练标签或模型微调,而是自动构建并优化针对数万对商品类别-属性组合的提示词。基于人工编写的初始提示,级联过程逐步调整指令以满足目录特定需求。该方法在大规模复杂工业目录中弥合了通用语言理解与领域知识之间的差距。大量实证评估显示,相较于传统的思维链提示,该方法在精确率和召回率上提升8%-10%。尤为关键的是,它将领域专家的工作量从每属性5.1小时减少至3分钟,降幅达99%。此外,该级联在五种语言及多种质量评估任务中均表现出良好泛化能力,持续保持性能优势。
原文摘要 · Abstract (English)
We introduce a novel, training free cascade for auto-prompting Large Language Models (LLMs) to assess product quality in e-commerce. Our system requires no training labels or model fine-tuning, instead automatically generating and refining prompts for evaluating attribute quality across tens of thousands of product category-attribute pairs. Starting from a seed of human-crafted prompts, the cascade progressively optimizes instructions to meet catalog-specific requirements. This approach bridges the gap between general language understanding and domain-specific knowledge at scale in complex industrial catalogs. Our extensive empirical evaluations shows the auto-prompt cascade improves precision and recall by $8-10\%$ over traditional chain-of-thought prompting. Notably, it achieves these gains while reducing domain expert effort from 5.1 hours to 3 minutes per attribute - a $99\%$ reduction. Additionally, the cascade generalizes effectively across five languages and multiple quality assessment tasks, consistently maintaining performance gains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。