构建电商相关性评估新基准,涵盖长尾难题与视觉显著性任务。
RAIR: A Rule-Aware Benchmark Uniting Challenging Long-Tail and Visual Salience Subset for E-commerce Relevance Assessment
- 提出规则感知的多子集评测框架,覆盖通用、长尾与视觉显著性场景。
- 14个模型测试显示,即使GPT-5在长尾和视觉任务上仍有明显短板。
- 专为中文电商场景设计,适合评估大模型与图文模型的真实能力。
搜索相关性在电商网站中至关重要。尽管大语言模型(LLMs)在相关性任务上表现优异,但现有基准缺乏足够复杂度,导致行业缺乏统一的评估标准。为此,我们提出了规则感知的相关性评估基准RAIR,基于真实场景的中文数据集。RAIR建立标准化评估框架,提供一套通用规则,奠定评估基础。同时,分析当前相关性模型所需的核心能力,构建包含三个子集的综合数据集:(1) 行业均衡采样的通用子集,用于评估基础能力;(2) 关注挑战性案例的长尾难例子集,评估性能极限;(3) 用于评估多模态理解的视觉显著性子集。我们在RAIR上对14个开源与闭源模型进行了实验,结果表明,即使是最优模型如GPT-5也面临显著挑战。RAIR数据现已公开,成为行业相关性评估的新基准,并为通用大模型与视觉语言模型(VLM)评估提供新视角。
原文摘要 · Abstract (English)
Search relevance plays a central role in web e-commerce. While large language models (LLMs) have shown significant results on relevance task, existing benchmarks lack sufficient complexity for comprehensive model assessment, resulting in an absence of standardized relevance evaluation metrics across the industry. To address this limitation, we propose Rule-Aware benchmark with Image for Relevance assessment(RAIR), a Chinese dataset derived from real-world scenarios. RAIR established a standardized framework for relevance assessment and provides a set of universal rules, which forms the foundation for standardized evaluation. Additionally, RAIR analyzes essential capabilities required for current relevance models and introduces a comprehensive dataset consists of three subset: (1) a general subset with industry-balanced sampling to evaluate fundamental model competencies; (2) a long-tail hard subset focus on challenging cases to assess performance limits; (3) a visual salience subset for evaluating multimodal understanding capabilities. We conducted experiments on RAIR using 14 open and closed-source models. The results demonstrate that RAIR presents sufficient challenges even for GPT-5, which achieved the best performance. RAIR data are now available, serving as an industry benchmark for relevance assessment while providing new insights into general LLM and Visual Language Model(VLM) evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。