用大模型自动评文本,还能解释并引用依据。
Rate, Explain and Cite (REC): Enhanced Explanation and Attribution in Automatic Evaluation by Large Language Models
- 用微调大模型从忠实性、指令遵循等维度评分
- 70B版本在多个基准上超越现有方法,解释更准确
- 适合需要可解释评估的生成内容质检场景
大语言模型在生成连贯高质量文本方面表现出色,但在内容质量评估方面仍面临挑战,如事实错误和幻觉问题。本文提出三种微调的通用大语言模型自动评估器:REC-8B、REC-12B 和 REC-70B,专门用于评估生成文本在忠实性、指令遵循、连贯性和完整性等方面的表现。这些模型不仅提供评分,还给出详细解释与可验证的引用,提升评估可信度。支持多种引用模式,适应不同延迟与粒度需求。在多个基准上的广泛评估表明,REC-70B 模型优于当前最先进的 LLM 评估方法,在内容评估中提供更高质量的解释与引用,且偏差最小。相关数据集与模型已开源:https://github.com/adelaidehsu/REC。
原文摘要 · Abstract (English)
LLMs have demonstrated impressive proficiency in generating coherent and high-quality text, making them valuable across a range of text-generation tasks. However, rigorous evaluation of this generated content is crucial, as ensuring its quality remains a significant challenge due to persistent issues such as factual inaccuracies and hallucination. This paper introduces three fine-tuned general-purpose LLM autoevaluators, REC-8B, REC-12B and REC-70B, specifically designed to evaluate generated text across several dimensions: faithfulness, instruction following, coherence, and completeness. These models not only provide ratings for these metrics but also offer detailed explanation and verifiable citation, thereby enhancing trust in the content. Moreover, the models support various citation modes, accommodating different requirements for latency and granularity. Extensive evaluations on diverse benchmarks demonstrate that our general-purpose LLM auto-evaluator, REC-70B, outperforms state-of-the-art LLMs, excelling in content evaluation by delivering better quality explanation and citation with minimal bias. Our REC dataset and models are available at https://github.com/adelaidehsu/REC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。