用多大模型迭代评审,自动评估自动生成问题的质量。
MIRROR: A Novel Approach for the Automated Evaluation of Open-Ended Question Generation
- 通过多大模型迭代反馈优化评分,替代人工评估。
- 评测相关性、恰当性等指标接近人类基准,相关性提升显著。
- 适合需要大规模自动评估问答质量的研究者和开发者。
自动问题生成是一项关键任务,需评估问题的吸引力、教学价值及激发批判性思维的能力。这些方面需要类人理解与判断,当前自动化系统尚不具备。然而,人工评估成本高,难以应用于大规模生成问题样本。为此,我们提出一种新系统 MIRROR(多大模型迭代评审与反馈优化评分),利用大语言模型(LLMs)自动化评估自动生成问题的质量。实验对比了 GPT-4、Gemini 及 Llama2-70b 等主流 LLM。结果显示,采用基于反馈的 MIRROR 方法后,相关性、恰当性、新颖性、复杂性和语法正确性等人类评估指标得分均提升,更接近人类基准。此外,与人类专家相比,GPT-4 在使用 MIRROR 时的皮尔逊相关系数也显著提高。错误分析表明,该方法显著改善了相关性与恰当性。
原文摘要 · Abstract (English)
Automatic question generation is a critical task that involves evaluating question quality by considering factors such as engagement, pedagogical value, and the ability to stimulate critical thinking. These aspects require human-like understanding and judgment, which automated systems currently lack. However, human evaluations are costly and impractical for large-scale samples of generated questions. Therefore, we propose a novel system, MIRROR (Multi-LLM Iterative Review and Response for Optimized Rating), which leverages large language models (LLMs) to automate the evaluation process for questions generated by automated question generation systems. We experimented with several state-of-the-art LLMs, such as GPT-4, Gemini, and Llama2-70b. We observed that the scores of human evaluation metrics, namely relevance, appropriateness, novelty, complexity, and grammaticality, improved when using the feedback-based approach called MIRROR, tending to be closer to the human baseline scores. Furthermore, we observed that Pearson's correlation coefficient between GPT-4 and human experts improved when using our proposed feedback-based approach, MIRROR, compared to direct prompting for evaluation. Error analysis shows that our proposed approach, MIRROR, significantly helps to improve relevance and appropriateness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。