用双检测器融合与选择性掩码提升中文拼写纠错效果
A Coin Has Two Sides: A Novel Detector-Corrector Framework for Chinese Spelling Correction
- 设计双检测器,分别优化精确率和召回率
- 通过特征融合与选择性掩码提升纠错准确率
- 适合需要高精度中文文本纠错的场景
中文拼写纠错(CSC)是自然语言处理的基础任务,旨在修正中文文本中的错别字。现有方法常采用独立的错误检测器定位错误位置,但检测器性能受限,导致精确率与召回率难以兼顾。本文提出一种新型检测-纠正框架,设计两个具有高精确率和高召回率的检测结果。考虑到错误具有上下文依赖性且检测结果可能不够精准,本文引入创新的特征融合策略和选择性掩码机制,将检测结果有效融入纠错过程。在主流CSC数据集上的实验验证了该方法的有效性。
原文摘要 · Abstract (English)
Chinese Spelling Correction (CSC) stands as a foundational Natural Language Processing (NLP) task, which primarily focuses on the correction of erroneous characters in Chinese texts. Certain existing methodologies opt to disentangle the error correction process, employing an additional error detector to pinpoint error positions. However, owing to the inherent performance limitations of error detector, precision and recall are like two sides of the coin which can not be both facing up simultaneously. Furthermore, it is also worth investigating how the error position information can be judiciously applied to assist the error correction. In this paper, we introduce a novel approach based on error detector-corrector framework. Our detector is designed to yield two error detection results, each characterized by high precision and recall. Given that the occurrence of errors is context-dependent and detection outcomes may be less precise, we incorporate the error detection results into the CSC task using an innovative feature fusion strategy and a selective masking strategy. Empirical experiments conducted on mainstream CSC datasets substantiate the efficacy of our proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。