用语言挖掘替代分类,高效构建法语克里奥尔语语料库
KréyoLID From Language Identification Towards Language Mining
- 将语言识别视为数据挖掘问题,聚焦低频语言
- 新流水线使语料构建速度更快、覆盖更广
- 适合资源有限的少数语言研究者
自动语言识别常被当作多分类任务处理。但在构建使用频率较低语言的数字语料库时,更应将其视为数据挖掘问题。对于这些语言,事先已知绝大多数文档价值有限。通过减少对无用文档的分类投入,可显著提升语料构建速度与覆盖范围。为验证语言挖掘视角的有效性,本文提出一种新流水线,并构建了多个基于法语的克里奥尔语语料库。
原文摘要 · Abstract (English)
Automatic language identification is frequently framed as a multi-class classification problem. However, when creating digital corpora for less commonly written languages, it may be more appropriate to consider it a data mining problem. For these varieties, one knows ahead of time that the vast majority of documents are of little interest. By minimizing resources spent on classifying such documents, we can create corpora much faster and with better coverage than using established pipelines. To demonstrate the effectiveness of the language mining perspective, we introduce a new pipeline and corpora for several French-based Creoles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。