arXiv:2509.26038cs.CL2025-09被引 2

用错误解释匹配例句,提升中文语法纠错准确率

RE$^2$: Improving Chinese Grammatical Error Correction via Retrieving Appropriate Examples with Explanation

  • 基于错误解释而非字面相似度检索纠错范例
  • 在两个数据集上显著提升纠错效果
  • 构建高质量中文语法错误解释数据集

中文语法纠错(CGEC)旨在识别并修正汉语句子中的语法错误。近期研究显示,大语言模型(LLMs)在该任务中已取得显著成果。对于LLM而言,选择合适的参考例句有助于提升性能。然而,现有方法主要依赖文本相似度进行例句检索,常因匹配实际错误模式不准确,导致检索到语义相近但语法无关的句子。为此,我们提出一种名为RE$^2$的方法,通过使用语法错误解释来检索恰当的参考例句,帮助LLM改进CGEC表现。我们在两个CGEC数据集上进行了实验,并构建了一个高质量的语法错误解释(GEE)数据集,不仅服务于本研究,也为未来CGEC与GEE研究提供宝贵资源。实验结果表明,所提方法能有效提升CGEC性能。

原文摘要 · Abstract (English)

The primary objective of Chinese grammatical error correction (CGEC) is to detect and correct errors in Chinese sentences. Recent research shows that large language models (LLMs) have been applied to CGEC with significant results. For LLMs, selecting appropriate reference examples can help improve their performance. However, existing methods predominantly rely on text similarity for example retrieval, a strategy that frequently mismatches actual error patterns and retrieves lexically similar yet grammatically irrelevant sentences. To address this problem, we propose a method named RE$^2$, which retrieves appropriate examples with explanations of grammatical errors. Instead of using text similarity of the input sentence, we use explanations of grammatical errors to select reference examples, which are used by LLMs to improve the performance of CGEC. We conduct experiments on two CGEC datasets and create a high-quality grammatical error explanation (GEE) dataset, which is not only used in our research but also serves as a valuable resource for future studies in both CGEC and GEE. The experimental results on the two datasets indicate that our proposed method effectively improves the performance of CGEC.

语法纠错大模型例句检索错误解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。