无需标注即可评估大模型生成质量,支持跨语言对比。
PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models
- 基于信息论设计无监督评分机制,不依赖人工标注。
- 在长篇问答任务中表现优于同类指标,与人类评分高度相关。
- 适合快速筛选最优大模型,尤其适用于缺乏标注数据的场景。
我们提出 PPLqa,一种简单可计算、语言无关、基于信息论的无监督度量方法,用于评估生成式大语言模型(LLMs)响应质量,无需真实答案或人工监督。该方法通过单一指标对大模型进行质量排序,帮助用户为特定任务选择最优模型。其评估涵盖连贯性与流畅性(写作质量)以及相关性与一致性(响应适切性),但并非显式依赖这些维度。PPLqa 在性能上媲美其他相关度量,在长篇问答任务中表现更优,且与人类及大模型评分具有良好相关性。该方法可避免传统评估所需的冗长标注流程。
原文摘要 · Abstract (English)
We propose PPLqa, an easy to compute, language independent, information-theoretic metric to measure the quality of responses of generative Large Language Models (LLMs) in an unsupervised way, without requiring ground truth annotations or human supervision. The method and metric enables users to rank generative language models for quality of responses, so as to make a selection of the best model for a given task. Our single metric assesses LLMs with an approach that subsumes, but is not explicitly based on, coherence and fluency (quality of writing) and relevance and consistency (appropriateness of response) to the query. PPLqa performs as well as other related metrics, and works better with long-form Q\&A. Thus, PPLqa enables bypassing the lengthy annotation process required for ground truth evaluations, and it also correlates well with human and LLM rankings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。