构建难识别语言数据集,测试系统在近亲语言和乱码情况下的表现
CHALIS: A Challenge Dataset for Language Identification in Difficult Scenarios
- 聚焦近亲语言与拼写噪声,构造双类挑战样本
- 多系统测试显示低资源语言与转写文本下准确率大幅下降
- 适合评估语言识别鲁棒性,尤其关注小语种场景
我们提出CHALIS(Challenging Language Identification Samples)——一个专为应对语言识别中难题设计的新基准数据集,涵盖近亲语言对(捷克/斯洛伐克、西班牙/加泰罗尼亚、葡萄牙/加利西亚、丹麦/挪威)及拼写噪声场景。通过跨文字转写、移除变音符号、模拟同形异义攻击和使用网络俚语,构建复杂输入。在该数据集上评估四个主流语言识别系统,发现其在低资源语言对和转写文本上表现显著下降。数据集已公开于https://huggingface.co/datasets/michal-tichy/CHALIS。
原文摘要 · Abstract (English)
We present CHALIS (Challenging Language Identification Samples), a new benchmark dataset explicitly designed to address difficult cases in language identification: cousin languages and orthographic noise. Our dataset has two parts: First, we collected sentences shared across mutually intelligible language pairs (Czech/Slovak, Spanish/Catalan, Portuguese/Galician, Danish/Norwegian). The second part tests for orthography noise: we transliterate text across multiple scripts, remove diacritics, simulate homoglyph attacks, and use Internet slang. We evaluate four widely used language identification systems on CHALIS and demonstrate that all struggle substantially in these scenarios, especially on lower-resource languages within cousin pairs and on transliterated input. The resource is publicly available at https://huggingface.co/datasets/michal-tichy/CHALIS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。