arXiv:2503.11858cs.CL2025-03被引 7

开源可解释的自然语言生成评估工具,支持多维度反馈。

OpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMs

  • 采用两阶段开源大模型集成,实现无参考文本评估。
  • 在多个任务上相关性优于现有方法,解释准确率超两倍。
  • 适合需要透明评估结果的研究者与开发者使用。

大型语言模型(LLMs)在自然语言生成(NLG)系统评估中展现出巨大潜力,可实现高质量、无需参考文本、多维度的评估。然而,现有基于LLM的评估指标存在两大缺陷:依赖专有模型生成训练数据或执行评估,缺乏细粒度且可解释的反馈。本文提出OpeNLGauge,一个完全开源、无需参考文本的NLG评估指标,能够基于错误片段提供精确解释。OpeNLGauge以两个大型开源模型的两阶段集成形式提供,或作为小型微调的评估模型,经验证具备对未见任务、领域和方面的泛化能力。大规模元评估表明,OpeNLGauge在相关性上达到竞争力水平,在某些任务上优于现有最先进模型,同时保证完全可复现,并提供解释精度超过两倍于现有方法的结果。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated great potential as evaluators of NLG systems, allowing for high-quality, reference-free, and multi-aspect assessments. However, existing LLM-based metrics suffer from two major drawbacks: reliance on proprietary models to generate training data or perform evaluations, and a lack of fine-grained, explanatory feedback. In this paper, we introduce OpeNLGauge, a fully open-source, reference-free NLG evaluation metric that provides accurate explanations based on error spans. OpeNLGauge is available as a two-stage ensemble of larger open-weight LLMs, or as a small fine-tuned evaluation model, with confirmed generalizability to unseen tasks, domains and aspects. Our extensive meta-evaluation shows that OpeNLGauge achieves competitive correlation with human judgments, outperforming state-of-the-art models on certain tasks while maintaining full reproducibility and providing explanations more than twice as accurate.

NLG评估可解释性开源模型大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。