arXiv:2607.07469cs.CLcs.AI2026-07

用多模型投票验证生成电商属性标签,低成本实现高精度数据标注。

SynthAVE: Scalable Synthetic Labeling for E-Commerce with LLM-Arena Validation

论文配图:SynthAVE: Scalable Synthetic Labeling for E-Commerce with LLM-Arena Validation
图 1 · 摘自论文原文
  • 通过21种不同模型配置的投票机制验证合成标签质量。
  • 最终标签准确率达97.9%,与专家标注一致性达Cohen's κ=0.92。
  • 适合需要大规模高质量标注的工业级电商场景。

为电商属性抽取微调大语言模型需涵盖数千种商品类型、属性及多语言的代表性标注数据,其组合规模可达数百万条,人工标注成本过高。尽管已有研究利用大模型生成合成标签,但在工业级部署中仍需集成质量控制机制。本文提出SynthAVE,一个覆盖12,726件商品、229种商品类型、792个属性及4种语言(西班牙语、法语、意大利语、德语)的大规模属性值验证基准。为实现大规模标签验证,我们引入多模型评估框架,每个样本由21种判别配置(7个模型族 × 3种提示)评估,最终标签通过多数投票决定;与合成标签不一致的样本由专家裁定,一致样本则分层抽样审计。多数投票集成结果与专家标注的Cohen's κ达到0.92(95.0%一致),而单个判别者间仅中等一致(Fleiss' κ=0.76)——此设计旨在保持多样性。结果表明,多样模型的个体判断可聚合为高度可靠的预测,实现低成本规模化验证,同时将专家精力集中于影响决策的分歧案例。我们估计最终标签质量为97.9%。

原文摘要 · Abstract (English)

Fine-tuning large language models (LLMs) for e-commerce attribute extraction requires labeled data representative across thousands of product types, attributes, and multiple languages. This combinatorial scale translates to millions of annotations, rendering human labeling prohibitively costly. While recent work has demonstrated synthetic label generation using LLMs, deploying such approaches at industrial scale requires integrated quality control mechanisms. We present SynthAVE, a large-scale benchmark for attribute-value verification spanning 12,726 products across 229 product types, 792 attributes, and 4 languages (Spanish, French, Italian, German). To validate synthetic labels at scale, we introduce a multi-LLM arena framework where each sample is evaluated by 21 judge configurations (7 model families $\times$ 3 prompts), with final labels determined via majority voting; disagreements with the synthetic label are expert-adjudicated and agreement cases are audited on a stratified sample. The majority vote ensemble agrees with expert labels at Cohen's $κ= 0.92$ (95.0% agreement), while individual judges agree with one another only moderately (Fleiss' $κ= 0.76$)--by design, since we select for diversity. This demonstrates that diverse models with varying individual judgments aggregate into highly reliable predictions, enabling cost-effective validation at scale while concentrating expert effort on the cases where it changes the label. We estimate the resulting label quality at 97.9%.

电商合成数据大模型标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。