提出双视角自动构建基准的NLG评估框架,提升可解释性。
A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better Interpretability
- 从双视角出发,区分不同评估能力,增强结果可解释性。
- 无需人工标注,自动构建评测基准,支持16个大模型测试。
- 适合关注评估可靠性和自动化评测的AI研究者。
在自然语言生成(NLG)元评估中,评估指标通常基于与人类评分的一致性进行检验。然而,传统方法存在处理人类评分不严谨、相关性度量选择模糊等问题,削弱了元评估的有效性。本文提出一种双视角NLG元评估框架,聚焦于不同评估能力,从而提升结果的可解释性。同时,提出一种无需新人工标注即可自动构建对应基准的方法。此外,基于该框架对16个代表性大语言模型(LLMs)进行了实验,从多个角度全面分析其评估表现。
原文摘要 · Abstract (English)
In NLG meta-evaluation, evaluation metrics are typically assessed based on their consistency with humans. However, we identify some limitations in traditional NLG meta-evaluation approaches, such as issues in handling human ratings and ambiguous selections of correlation measures, which undermine the effectiveness of meta-evaluation. In this work, we propose a dual-perspective NLG meta-evaluation framework that focuses on different evaluation capabilities, thereby providing better interpretability. In addition, we introduce a method of automatically constructing the corresponding benchmarks without requiring new human annotations. Furthermore, we conduct experiments with 16 representative LLMs as the evaluators based on our proposed framework, comprehensively analyzing their evaluation performance from different perspectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。