清理维基百科低质内容,提升多语言模型训练效果
How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP
- 对非英语维基进行噪声过滤,移除大量低质数据
- 过滤后模型性能提升,尤其在低质量语种上收益显著
- 提出四级质量分级体系,适合多语言NLP研究者参考
维基百科因其广泛的语言覆盖和高可信度被视为NLP关键资源。然而在低资源与多语言场景中,其质量受到质疑。本文对非英语维基进行全面过滤,采用通常用于网络文本去噪的方法,移除了大量数据。分析发现存在脚本混杂、语言污染、重复模板文章及机器人生成内容集中等问题。据此构建了四级质量评级体系,与其它质量评估方法高度一致。在三个语言建模任务中验证,使用过滤后数据训练的模型性能普遍优于或匹配原始维基数据,低质量语种版块提升最明显。本研究为维基百科在NLP中的高质量使用提供初步实践指南,推动未来数据集创建与清洗标准建设。
原文摘要 · Abstract (English)
Wikipedia's perceived high quality and broad language coverage have established it as a fundamental resource in NLP. However, in recent years, such assumptions of high quality have become the subject of scrutiny in low-resource and multilingual contexts. In this study, we subject the entirety of non-English Wikipedia to a data filtering procedure typically reserved for noisy web-text -- a process which removes a large percentage of the collection's data. In analysing the removed data, we reveal numerous systematic quality issues, such as script and language contamination, repeated template and placeholder articles, and a high concentration of bot-generated content. We consolidate these findings into a 4-level quality ranking of Wikipedia, which shows strong correspondence with alternative quality measures and heuristics. Lastly, we evaluate the downstream impact of quality filtering in three practical language modelling scenarios, showing that models trained on filtered data largely match or outperform those trained on raw Wikipedia, with the largest gains observed for lower-quality language editions. Ultimately, our experiments serve as a first step in establishing quality-aware best practices for Wikipedia utilization in NLP, laying groundwork that can inform future dataset creation and curation efforts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。