arXiv:2412.12832cs.CLcs.AI2024-12AAAI被引 2

用动态权重提升语法纠错模型评估准确性

DSGram: Dynamic Weighting Sub-Metrics for Grammatical Error Correction in the Era of Large Language Models

  • 融合语义连贯、编辑层级和流畅性,动态分配评估权重
  • 在多个数据集上显著提升评估相关性,优于传统方法
  • 适合需要可靠评估的GEC研究与开发人员

语法纠错(GEC)模型的评估日益困难,因大语言模型(LLM)生成的修正结果常偏离标准参考答案,削弱了传统基于参考的评估指标可靠性。本文提出一种新评估框架DSGram,整合语义连贯性、编辑层级和流畅性,并引入动态权重机制。该框架结合分析层次过程(AHP)与大语言模型,确定各评估维度的相对重要性。此外,我们构建了一个包含人工标注与LLM模拟句子的数据集,用于验证算法并微调更低成本的模型。实验表明,所提方法显著提升了GEC模型评估的有效性。

原文摘要 · Abstract (English)

Evaluating the performance of Grammatical Error Correction (GEC) models has become increasingly challenging, as large language model (LLM)-based GEC systems often produce corrections that diverge from provided gold references. This discrepancy undermines the reliability of traditional reference-based evaluation metrics. In this study, we propose a novel evaluation framework for GEC models, DSGram, integrating Semantic Coherence, Edit Level, and Fluency, and utilizing a dynamic weighting mechanism. Our framework employs the Analytic Hierarchy Process (AHP) in conjunction with large language models to ascertain the relative importance of various evaluation criteria. Additionally, we develop a dataset incorporating human annotations and LLM-simulated sentences to validate our algorithms and fine-tune more cost-effective models. Experimental results indicate that our proposed approach enhances the effectiveness of GEC model evaluations.

语法纠错评估方法大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。