arXiv:2511.17562cs.CLcs.AI2025-11被引 2

中文错别字与语法纠错新模型,性能领先

ChineseErrorCorrector3-4B: State-of-the-Art Chinese Spelling and Grammar Corrector

  • 基于Qwen3-4B构建统一纠错模型
  • 在多个基准数据集上F1和F0.5均夺冠
  • 适合中文文本校对与自然语言处理研究

本文介绍ChineseErrorCorrector3-4B,一个基于Qwen3-4B的中文统一拼写与语法纠错模型。该模型在通用文本纠错任务中表现优异,在拼写纠错(CSC)和语法纠错(CGC)任务中均达到当前最佳水平。在SIGHAN-2015、EC-LAW、MCSC和NaCGEC等多个权威基准数据集上,其F1和F0.5分数显著超越现有公开模型,两项任务均排名第一。

原文摘要 · Abstract (English)

This paper introduces ChineseErrorCorrector3-4B, a unified model for Chinese spelling and grammatical error correction based on Qwen3-4B. The model demonstrates outstanding performance in general text correction tasks and achieves state-of-the-art results in both spelling correction (CSC) and grammatical correction (CGC). On several authoritative benchmark datasets -- including SIGHAN-2015, EC-LAW, MCSC, and NaCGEC -- the model's F1 and F0.5 scores significantly surpass existing publicly available models, ranking first in both spelling and grammatical error correction tasks.

中文纠错大模型语法纠错Qwen

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。