审计伦巴第语语料库,发现数据量虚高且存在严重偏见。
"Chi nas dal soch el sent de legn" -- Auditing Text Corpora for Lombard
- 人工审核伦巴第语语料,发现网页爬取数据多含错误识别和噪声。
- 高质量数据集中西部伦巴第语占主导,东部方言严重缺失。
- 呼吁以社区参与方式构建多样化的语言数据集。
全球许多语言在自然语言处理工具方面仍属资源匮乏,主要因缺乏高质量数据集用于训练、开发和评估各类任务(如机器翻译)。本文对意大利伦巴第语(一种资源匮乏的语言连续体)的平行语料与单语语料进行了人工审计。分析显示,看似丰富的网络爬取数据实为幻觉:大规模数据集充斥着严重的语言误标、模板文本及非语言噪声。此外,我们分析了各类语料中有效伦巴第语部分的拼写构成,发现所有语料均存在拼写系统冲突与严重代表性偏差:高质量数据高度偏向西部伦巴第方言,东部方言几乎被忽略。这一结果凸显了亟需以方言意识为导向、由社区驱动的数据整理,而非单纯追求数据量的自动采集。
原文摘要 · Abstract (English)
Several of the world's languages are still under-resourced in terms of Natural Language Processing (NLP) tools. This is mostly due to the lack of high-quality datasets to train, develop, and evaluate systems and models for several tasks, such as Machine Translation (MT). We conduct a manual audit of the parallel and monolingual corpora available for Lombard, an under-resourced language continuum from Italy. Our analysis reveals that the perceived abundance of web-scraped data is an illusion, with massive datasets plagued by severe language misidentification, boilerplate text, and non-linguistic noise. Furthermore, we analyze the orthographic composition of the valid Lombard portions across web-scraped datasets, curated corpora, and benchmarks. Our findings show conflicting orthographical systems and severe representational bias across all corpora: high-quality data is heavily skewed towards Western Lombard varieties, with Eastern ones left on the margins. This underscores the need for variety-aware, community-driven data curation rather than purely quantity-driven scraping.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。