arXiv:2410.04292cs.CL2024-10被引 2

仅用20个样本就能发现多语言数据中的低质语种子集

Efficiently Identifying Low-Quality Language Subsets in Multilingual Datasets: A Case Study on a Large-Scale Multilingual Audio Dataset

  • 通过偏好比例检验法,仅标注20样本即可识别不可靠语种
  • 过滤低质数据后,跨语言语音转录准确率提升25.7%
  • 适合需要高效清洗多语言音频数据的研究者使用

构建多语言数据集极具挑战性。为提升可扩展性,研究者常引入不完美的分类器(如语言识别模型),但这些模型易出错,导致部分语言子集不可靠。本文提出一种统计检验方法——偏好比例检验,用于识别此类不可靠子集。通过对一个近期大规模多语言语音数据集X-IPAPack(Zhu et al., 2024)中10个语种子集各标注20个样本,我们成功识别出系统性转录错误。在下游语音音素转录任务中,剔除这些低质数据后,对分布外语言的转录性能显著提升,相对改进达25.7%。该方法为多语言数据集的系统性审计提供了可靠路径。

原文摘要 · Abstract (English)

Curating datasets that span multiple languages is challenging. To make the collection more scalable, researchers often incorporate one or more imperfect classifiers in the process, like language identification models. These models, however, are prone to failure, resulting in some language subsets being unreliable for downstream tasks. We introduce a statistical test, the Preference Proportion Test, for identifying such unreliable subsets. By annotating only 20 samples for a language subset, we're able to identify systematic transcription errors for 10 language subsets in a recent large multilingual transcribed audio dataset, X-IPAPack (Zhu et al., 2024). We find that filtering this low-quality data out when training models for the downstream task of phonetic transcription brings substantial benefits, most notably a 25.7% relative improvement on transcribing recordings in out-of-distribution languages. Our method lays a path forward for systematic and reliable multilingual dataset auditing.

多语言数据数据清洗语音转录统计检验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。