arXiv:2506.22402cs.CL2025-06中稿 · TSD 2025

用动态生成错误提升捷克语语法纠错模型性能。

Refining Czech GEC: Insights from a Multi-Experiment Approach

  • 通过实时合成错误数据增强训练,融合通用与捷克语特有错误。
  • 在多个实验中验证了数据规模、分词粒度等对效果的影响。
  • 模型兼具高精度与低计算开销,适合实际部署使用。

我们提出一个捷克语语法错误修正(GEC)系统,在该语言上达到当前最佳性能。系统基于Transformer架构的神经网络翻译方法,核心是实时合成错误的数据增强管道,能动态地向句子中引入语言无关和捷克语特有错误。我们开展了一系列全面实验,研究了捷克语GEC语料库作为错误注入基础的效果、多种错误生成策略、领域平衡、分词粒度、模型大小及微调阶段的数据缩放。此外,还在终端用户和专家微调场景下评估了大语言模型(LLMs)在捷克语GEC上的表现。最优模型在性能和计算效率上均更优。源代码与训练模型已公开于 https://github.com/ufal/tsd2025-gec。

原文摘要 · Abstract (English)

We present a grammar error correction (GEC) system that achieves state of the art for the Czech language. Our system is based on a neural network translation approach with the Transformer architecture, and its key feature is its real-time synthetic generation pipeline, which dynamically augments sentences with artificial errors by introducing both language-agnostic and Czech-specific errors. We conduct a comprehensive series of experiments, investigating the Czech GEC corpora as bases for synthetic error introduction, several error generation strategies, domain balancing, tokenization granularity, model size, and data scaling during fine-tuning. Additionally, we evaluate the performance of large language models (LLMs) on Czech GEC in both end-user and expert fine-tuning scenarios. Our best-performing model is superior both in performance and computational efficiency. The source code and the trained model links are available on https://github.com/ufal/tsd2025-gec.

语法纠错捷克语数据增强Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。