测试UMBRELA在不同大模型上的表现,发现其效果因模型大小而异。
Does UMBRELA Work on Other LLMs?
- 在多种大模型上复现UMBRELA评估框架,检验其泛化能力
- DeepSeek V3表现接近GPT-4o,LLaMA-3.3-70B略低,小模型更差
- 适合关注大模型评估一致性的研究者和开发者参考
我们在一系列大语言模型(LLMs)上复现了UMBRELA LLM Judge评估框架,以评估其在原研究之外的通用性。研究重点关注大模型选择对相关性评估准确率的影响,使用排行榜相关性与逐标签一致性指标进行评估。结果表明,使用DeepSeek V3时,UMBRELA表现与原始研究中使用的GPT-4o相当;对于LLaMA-3.3-70B,性能略低,且随着模型规模减小,性能进一步下降。
原文摘要 · Abstract (English)
We reproduce the UMBRELA LLM Judge evaluation framework across a range of large language models (LLMs) to assess its generalizability beyond the original study. Our investigation evaluates how LLM choice affects relevance assessment accuracy, focusing on leaderboard rank correlation and per-label agreement metrics. Results demonstrate that UMBRELA with DeepSeek V3 obtains very comparable performance to GPT-4o (used in original work). For LLaMA-3.3-70B we obtain slightly lower performance, which further degrades with smaller LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。