arXiv:2511.01090cs.CL2025-11被引 2

用多维度过滤提升罗马尼亚语大模型训练数据质量

Improving Romanian LLM Pretraining Data using Diversity and Quality Filtering

  • 构建轻量级多任务模型,对罗马尼亚语文本进行教育价值、主题、格式等多层筛选
  • 过滤后数据使大模型在多个基准测试中表现显著提升,罗马尼亚语与英语数据主题差异明显
  • 适合关注低资源语言大模型训练的研究者和开发者

大型语言模型(LLMs)近年来广受欢迎,许多任务上已达到甚至超越人类水平。其成功很大程度依赖于高质量训练数据的可用性与精心筛选。对于低资源语言而言,高质量语料库尤为稀缺,数据质量至关重要。本文研究了罗马尼亚语预训练语料的特征与覆盖范围,并分析其与英语数据的差异。通过在人工标注的罗马尼亚语文本上训练一个轻量级多任务模型,我们实现了多层次过滤(如教育价值、主题、格式),生成高质量的预训练数据集。实验表明,罗马尼亚语与英语数据在主题分布上存在显著差异,且经过过滤的数据显著提升了大模型在多个基准测试中的表现。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have recently exploded in popularity, often matching or outperforming human abilities on many tasks. One of the key factors in training LLMs is the availability and curation of high-quality data. Data quality is especially crucial for under-represented languages, where high-quality corpora are scarce. In this work we study the characteristics and coverage of Romanian pretraining corpora and we examine how they differ from English data. By training a lightweight multitask model on carefully LLM-annotated Romanian texts, we are able to analyze and perform multi-level filtering (e.g., educational value, topic, format) to generate high-quality pretraining datasets. Our experiments show noteworthy trends in the topics present in Romanian and English data, while also proving the effectiveness of filtering data through improved LLM pretraining performance across multiple benchmarks.

大模型低资源语言数据过滤

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。