用规则强化学习提升大模型语法纠错能力,效果优于传统方法。
Harnessing Rule-Based Reinforcement Learning for Enhanced Grammatical Error Correction
- 基于规则的强化学习框架,引导大模型更精准纠正语法错误。
- 在中文数据集上达到当前最优,召回率显著提升。
- 适合需要高准确性和可解释性的语法纠错场景。
语法纠错是自然语言处理中的重要任务。传统基于编码器-解码器的模型虽取得一定成效,但大模型在此领域的应用仍不充分。现有研究多依赖监督微调,直接让大模型生成修正句,限制了其推理能力。为此,我们提出一种基于规则的强化学习新框架。在中文数据集上的实验表明,该框架实现了当前最优性能,召回率明显提升。结果清晰展示了利用强化学习引导大模型的优势,为未来语法纠错提供了更可控、更可靠的范式。
原文摘要 · Abstract (English)
Grammatical error correction is a significant task in NLP. Traditional methods based on encoder-decoder models have achieved certain success, but the application of LLMs in this field is still underexplored. Current research predominantly relies on supervised fine-tuning to train LLMs to directly generate the corrected sentence, which limits the model's powerful reasoning ability. To address this limitation, we propose a novel framework based on Rule-Based RL. Through experiments on the Chinese datasets, our Rule-Based RL framework achieves \textbf{state-of-the-art }performance, with a notable increase in \textbf{recall}. This result clearly highlights the advantages of using RL to steer LLMs, offering a more controllable and reliable paradigm for future development in GEC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。