arXiv:2604.19593cs.CLcs.AI2026-04

首个罗马尼亚语法律领域语法纠错数据集,助力法律文本精准校对。

RoLegalGEC: Legal Domain Grammatical Error Detection and Correction Dataset for Romanian

论文配图:RoLegalGEC: Legal Domain Grammatical Error Detection and Correction Dataset for Romanian
图 1 · 摘自论文原文
  • 构建了含35万条法律文本错误的平行数据集
  • 在该数据集上验证了多种神经网络模型的纠错效果
  • 适合法律AI、自然语言处理研究者使用

法律文件的清晰准确至关重要,因此面向法律专业人士的语法纠错工具必须理解法律语境中的错误并进行相应修正,这要求模型需在真实法律数据上训练。然而,类似罗马尼亚语这类语言的法律领域人工标注数据极为稀缺。现有常用方法为合成平行数据,但依赖对罗马尼亚语语法的结构化理解。本文首次提出罗法律语法纠错数据集(RoLegalGEC),包含35万条法律文本中的语法错误及其标注。我们评估了多种神经网络模型,包括知识蒸馏的Transformer、用于检测的序列标注架构,以及多种预训练文本到文本Transformer模型用于纠错。该数据集与模型组合将丰富罗马尼亚语法律NLP研究资源。

原文摘要 · Abstract (English)

The importance of clear and correct text in legal documents cannot be understated, and, consequently, a grammatical error correction tool meant to assist a professional in the law must have the ability to understand the possible errors in the context of a legal environment, correcting them accordingly, and implicitly needs to be trained in the same environment, using realistic legal data. However, the manually annotated data required by such a process is in short supply for languages such as Romanian, much less for a niche domain. The most common approach is the synthetic generation of parallel data; however, it requires a structured understanding of the Romanian grammar. In this paper, we introduce, to our knowledge, the first Romanian-language parallel dataset for the detection and correction of grammatical errors in the legal domain, RoLegalGEC, which aggregates 350,000 examples of errors in legal passages, along with error annotations. Moreover, we evaluate several neural network models that transform the dataset into a valuable tool for both detecting and correcting grammatical errors, including knowledge-distillation Transformers, sequence tagging architectures for detection, and a variety of pre-trained text-to-text Transformer models for correction. We consider that the set of models, together with the novel RoLegalGEC dataset, will enrich the resource base for further research on Romanian.

语法纠错法律NLP数据集罗马尼亚语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。