arXiv:2412.13268cs.IR2024-12被引 8

用多个小模型合评搜索相关性,省钱又准。

JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment

  • 融合多个小模型或提示的判断结果
  • 在基准测试中表现接近顶级大模型
  • 适合预算有限但需可靠评估的研究者

有效训练和评估检索系统需要大量相关性判断,传统上依赖人工标注,成本高且耗时。大型语言模型(LLMs)在生成搜索任务相关性标签方面展现出潜力,可作为人工评估的替代方案。当前方法多依赖单一大模型(如GPT-4),虽有效但昂贵,且存在模型内偏见,可能偏向使用相似模型的系统。本文提出JudgeBlender框架,利用多个小型开源模型,通过组合多模型(LLMBlender)或多提示(PromptBlender)的评估结果来生成相关性判断。基于LLMJudge基准[18],我们对比了JudgeBlender与现有先进方法及LLMJudge挑战赛优胜者。结果显示,JudgeBlender实现了具有竞争力的性能,表明在可靠相关性评估中,超大规模模型往往并非必需。

原文摘要 · Abstract (English)

The effective training and evaluation of retrieval systems require a substantial amount of relevance judgments, which are traditionally collected from human assessors -- a process that is both costly and time-consuming. Large Language Models (LLMs) have shown promise in generating relevance labels for search tasks, offering a potential alternative to manual assessments. Current approaches often rely on a single LLM, such as GPT-4, which, despite being effective, are expensive and prone to intra-model biases that can favour systems leveraging similar models. In this work, we introduce JudgeBlender, a framework that employs smaller, open-source models to provide relevance judgments by combining evaluations across multiple LLMs (LLMBlender) or multiple prompts (PromptBlender). By leveraging the LLMJudge benchmark [18], we compare JudgeBlender with state-of-the-art methods and the top performers in the LLMJudge challenge. Our results show that JudgeBlender achieves competitive performance, demonstrating that very large models are often unnecessary for reliable relevance assessments.

信息检索LLM评估模型集成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。