用大模型自动评估推荐系统,无需人工参与。
No-Human in the Loop: Agentic Evaluation at Scale for Recommendation
- 构建多智能体框架,通过多数投票生成可靠标签。
- GPT-4o在延迟、准确率和成本上表现最佳。
- 适合需要大规模自动化评估的推荐系统研究者。
评估大语言模型作为评判者的可靠性对构建可扩展、可信的评估流程至关重要。我们提出ScalingEval,一项大规模基准测试,系统比较了36个大模型(包括GPT、Gemini、Claude和Llama)在多个产品类别中的表现,采用共识驱动的评估协议。我们的多智能体框架通过可扩展的多数投票机制,将模式审计和问题代码聚合为真实标签,实现无需人工标注的可复现评估。应用于大规模互补商品推荐任务,该基准报告了四项关键发现:(i) Anthropic Claude 3.5 Sonnet决策置信度最高;(ii) Gemini 1.5 Pro在各类别中综合表现最优;(iii) GPT-4o提供最佳延迟-精度-成本权衡;(iv) GPT-OSS 20B是开源模型中领先者。类别级分析显示,结构化领域(电子、体育)达成强共识,而生活方式类(服饰、食品)仍存在显著分歧。这些结果确立了ScalingEval作为大模型作为评判者可复现的基准与评估协议,为规模化、可靠性及模型家族权衡提供行动指导。
原文摘要 · Abstract (English)
Evaluating large language models (LLMs) as judges is increasingly critical for building scalable and trustworthy evaluation pipelines. We present ScalingEval, a large-scale benchmarking study that systematically compares 36 LLMs, including GPT, Gemini, Claude, and Llama, across multiple product categories using a consensus-driven evaluation protocol. Our multi-agent framework aggregates pattern audits and issue codes into ground-truth labels via scalable majority voting, enabling reproducible comparison of LLM evaluators without human annotation. Applied to large-scale complementary-item recommendation, the benchmark reports four key findings: (i) Anthropic Claude 3.5 Sonnet achieves the highest decision confidence; (ii) Gemini 1.5 Pro offers the best overall performance across categories; (iii) GPT-4o provides the most favorable latency-accuracy-cost tradeoff; and (iv) GPT-OSS 20B leads among open-source models. Category-level analysis shows strong consensus in structured domains (Electronics, Sports) but persistent disagreement in lifestyle categories (Clothing, Food). These results establish ScalingEval as a reproducible benchmark and evaluation protocol for LLMs as judges, with actionable guidance on scaling, reliability, and model family tradeoffs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。