arXiv:2505.04507cs.CL2025-05

自动检测俄语诗歌中的拼写语法错误,提升生成模型训练数据质量

Detecting Spelling and Grammatical Anomalies in Russian Poetry Texts

  • 用无监督与有监督方法对比检测文本异常
  • 构建了首个俄语诗歌语法错误标注数据集RUPOR
  • 适合从事创意生成模型数据清洗的研究者使用

生成模型在诗歌等创造性任务中的表现高度依赖于微调数据集的语言质量。生成诗中出现的流畅性缺陷会显著降低其价值。然而,训练数据常来自缺乏严格质量控制的网络平台,给数据工程师带来管理缺陷水平的挑战。为此,我们提出使用自动化语言异常检测技术,识别并过滤低质量文本。本文全面比较了无监督与有监督的文本异常检测方法,采用合成数据与人工标注数据进行评估,并引入RUPOR数据集——一个用于跨句语法错误检测的俄语诗歌人工标注集合,同时提供完整评估代码。本工作旨在为社区提供工具与洞见,以提升创意领域生成模型训练数据的质量。

原文摘要 · Abstract (English)

The quality of natural language texts in fine-tuning datasets plays a critical role in the performance of generative models, particularly in computational creativity tasks such as poem or song lyric generation. Fluency defects in generated poems significantly reduce their value. However, training texts are often sourced from internet-based platforms without stringent quality control, posing a challenge for data engineers to manage defect levels effectively. To address this issue, we propose the use of automated linguistic anomaly detection to identify and filter out low-quality texts from training datasets for creative models. In this paper, we present a comprehensive comparison of unsupervised and supervised text anomaly detection approaches, utilizing both synthetic and human-labeled datasets. We also introduce the RUPOR dataset, a collection of Russian-language human-labeled poems designed for cross-sentence grammatical error detection, and provide the full evaluation code. Our work aims to empower the community with tools and insights to improve the quality of training datasets for generative models in creative domains.

自然语言处理数据清洗诗歌生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。