评估方法比模型选择更重要,决定属性抽取效果的关键是评价方式和数据质量。
Beyond Exact Match: How Evaluation Methodology Dominates Model Choice in LLM-Based Product Attribute Extraction

- 对比四种提示策略,在相同模型上测试评估方法影响更大。
- 评估方式导致的性能差异是模型选择的23倍,数据噪声率达23.2%。
- 适合做电商属性抽取的工程团队,应优先优化评估流程与标签质量。
大型语言模型(LLMs)已成为电商场景中结构化商品属性抽取的默认方案,但不同模型、数据集和提示策略间表现差异显著。本文在MAVE基准上,对GPT-4o-mini和Gemini 2.5 Flash两个生产级模型,采用零样本、少样本、模式引导和定义增强四种提示策略,共评估6,400次属性预测结果。使用精确匹配与模糊字符串匹配两种评价方式,并对真实标签进行严格噪声审计。我们正式分解了F1分数的方差来源,发现评估方法带来的波动约为模型选择的23倍、提示策略的5倍。进一步发现MAVE基准的真实标签噪声率为23.2%,针对现代LLM输出。成对置换检验(B=10,000)显示协议间F1差距极显著(p<0.0001),协议间一致性(Cohen's kappa)达0.769。结论:在实际应用中,评估方法与数据质量对属性抽取效果的影响远超模型选型与提示工程。
原文摘要 · Abstract (English)
Large language models (LLMs) have become a default choice for structured product attribute extraction in e-commerce pipelines, with practitioners reporting widely varying performance across models, datasets, and prompting strategies. This paper presents a controlled empirical study comparing four prompting strategies -- zero-shot, few-shot, schema-guided, and definition-augmented -- across two production-grade LLMs (GPT-4o-mini and Gemini 2.5 Flash) on the MAVE benchmark. We evaluate 6,400 attribute-level predictions using both exact and fuzzy string matching, and conduct a rigorous noise audit of the ground truth labels. We formally decompose F1 variance across four experimental factors and find that evaluation methodology produces variance approximately 23 times larger than model choice and 5 times larger than prompting strategy choice. We further establish that the MAVE benchmark exhibits a 23.2% ground truth noise rate against modern LLM outputs. Paired permutation tests (B=10,000) confirm that the inter-protocol F1 gap is highly significant (p<0.0001) and Cohen's kappa of 0.769 between protocols indicates substantial agreement. We conclude that for production attribute extraction pipelines, evaluation methodology and data quality dominate the impact of model selection and prompt engineering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。