用对比学习提升文本生成评估质量与效率
ContrastScore: Towards Higher Quality, Less Biased, More Efficient Evaluation Metrics with Contrastive Evaluation
- 基于对比学习设计新评估指标,减少对参考文本依赖
- 在翻译和摘要任务中相关性超越主流基线,3B模型胜过7B
- 有效缓解长度和概率偏好等常见评估偏差,适合实际应用
自动评估生成文本质量仍是重大挑战。传统基于参考文本的度量与人类评价相关性较弱。近期研究主张使用大语言模型(LLM)作为源文本度量进行自然语言生成(NLG)评估。尽管前景可观,但小规模模型仍难以准确匹配人类判断。本文提出ContrastScore,一种基于对比学习的评估方法,旨在实现更高品质、更少偏见、更高效的文本评估。我们在机器翻译和摘要两个任务上评估ContrastScore。实验表明,ContrastScore在各项任务中均显著优于单模型与集成基线,且基于Qwen 3B和0.5B的版本在相关性上甚至超过Qwen 7B,参数仅为其一半,展现出优异效率。此外,该方法有效缓解了长度偏好和似然偏好等常见评估偏差,提升了自动评估的鲁棒性。
原文摘要 · Abstract (English)
Evaluating the quality of generated text automatically remains a significant challenge. Conventional reference-based metrics have been shown to exhibit relatively weak correlation with human evaluations. Recent research advocates the use of large language models (LLMs) as source-based metrics for natural language generation (NLG) assessment. While promising, LLM-based metrics, particularly those using smaller models, still fall short in aligning with human judgments. In this work, we introduce ContrastScore, a contrastive evaluation metric designed to enable higher-quality, less biased, and more efficient assessment of generated text. We evaluate ContrastScore on two NLG tasks: machine translation and summarization. Experimental results show that ContrastScore consistently achieves stronger correlation with human judgments than both single-model and ensemble-based baselines. Notably, ContrastScore based on Qwen 3B and 0.5B even outperforms Qwen 7B, despite having only half as many parameters, demonstrating its efficiency. Furthermore, it effectively mitigates common evaluation biases such as length and likelihood preferences, resulting in more robust automatic evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。