用轻量投票机制让大模型自动评估自由问答,又准又省资源。
CLEV: LLM-Based Evaluation Through Lightweight Efficient Voting for Free-Form Question-Answering
- 两模型先评,意见不一致时才调第三模型,减少计算开销。
- 实验显示评估结果稳定,与人工评分高度一致,适合大规模测试。
- 适合需要高效、可靠评估大模型回答质量的研究者使用。
自由形式问答(Free-form QA)的评估因答案多样且开放而面临挑战。传统自动评估指标难以捕捉语义等价性或适应开放回答的多样性。利用大语言模型(LLMs)作为评估者,凭借其强大的语言理解与指令遵循能力,成为有前景的替代方案。本文提出共识轻量高效投票(CLEV)方法,采用两个主判别模型进行评估,并仅在二者意见不一致时调用第三个模型介入。该设计在保证评估可靠性的同时,显著降低不必要的计算成本。通过实验(包括人工评估)验证,CLEV能提供一致、可扩展且资源高效的评估,为大模型在自由问答任务上的评估提供稳健框架。
原文摘要 · Abstract (English)
Evaluating free-form Question Answering (QA) remains a challenge due to its diverse and open-ended nature. Traditional automatic metrics fail to capture semantic equivalence or accommodate the variability of open-ended responses. Leveraging Large Language Models (LLMs) as evaluators offers a promising alternative due to their strong language understanding and instruction-following capabilities. We propose Consensus via Lightweight Efficient Voting (CLEV), which employs two primary LLMs as judges and invokes a third judge only in cases of disagreement. This approach prioritizes evaluation reliability while reducing unnecessary computational demands. Through experiments, including human evaluation, we demonstrate CLEV's ability to provide consistent, scalable, and resource-efficient assessments, establishing it as a robust framework for evaluating LLMs on free-form QA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。