arXiv:2602.23845cs.CL2026-02

提出中文段落级语言与事实错误联合纠错新任务

CLFEC: A New Task for Unified Linguistic and Factual Error Correction in paragraph-level Chinese Professional Writing

  • 构建跨多领域的中文专业写作混合数据集
  • 联合纠错比分离处理效果更好,代理工作流有效
  • 适合中文自然语言处理与智能校对研究者

中文文本纠错传统上聚焦于拼写和语法,而事实性错误常被单独处理。但在段落级中文专业写作中,语言(词汇/语法/标点)与事实错误频繁共现且相互影响,且许多草稿级错误在编辑后发表文本中难以察觉,因此统一纠错既必要又需建立可控基准。本文提出CLFEC(中文语言与事实错误纠正)新任务,构建涵盖时政、金融、法律、医学的混合多领域中文专业写作数据集。系统研究基于大模型的纠错范式,从提示工程到检索增强生成(RAG)及代理工作流。分析发现实际挑战包括专用纠错模型泛化能力有限、事实修复需证据支撑、混合错误段落难处理,以及对无错输入的过度修正。结果表明,在同一上下文中联合处理语言与事实错误优于分离流水线,且合适的骨干模型下代理工作流可有效运行。总体而言,CLFEC为中文文本纠错研究提供新基准,并为校对系统设计提供实践指导。

原文摘要 · Abstract (English)

Chinese text correction has traditionally focused on spelling and grammar, while factual error correction is usually treated separately. However, in paragraph-level Chinese professional writing, linguistic (word/grammar/punctuation) and factual errors frequently co-occur and interact, while many draft-level errors are sparsely observable in published texts after editorial review, making unified correction both necessary and controlled benchmark construction essential. This paper introduces CLFEC (Chinese Linguistic \& Factual Error Correction), a new task for joint linguistic and factual correction. We construct a mixed, multi-domain Chinese professional writing dataset spanning current affairs, finance, law, and medicine. We then conduct a systematic study of LLM-based correction paradigms, from prompting to retrieval-augmented generation (RAG) and agentic workflows. The analysis reveals practical challenges, including limited generalization of specialized correction models, the need for evidence grounding for factual repair, the difficulty of mixed-error paragraphs, and over-correction on clean inputs. Results further show that handling linguistic and factual errors within the same context outperforms decoupled pipelines, and that agentic workflows can be effective with suitable backbone models. Overall, CLFEC provides a new benchmark for Chinese text correction research and practical guidance for proofreading systems.

文本纠错中文NLP大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。