arXiv:2507.11832cs.CL2025-07被引 1

针对印度多语言共用文字的难题,构建25万句语料库并提出高效识别模型。

ILID: Native Script Language Identification for Indian Languages

  • 基于23种印度语言构建大规模标注数据集,含22种官方语言
  • 在短文本与混合语境下实现超出现有模型的识别准确率
  • 适合多语言处理、低资源语言研究者使用

语言识别是自然语言处理的关键基础任务,常作为多语言机器翻译、信息检索、问答系统和文本摘要等应用的预处理步骤。其核心挑战在于区分噪声大、长度短且存在代码混杂的文本。这一问题在印度语言中尤为突出,因多种语言具有相似的词汇与发音特征,且多数共享同一书写系统,导致识别难度剧增。为此,本文构建并发布一个包含25万条句子的数据集,涵盖23种语言(包括英语及全部22种印度官方语言),其中多数语言的数据为全新生成。同时,我们开发并公开了基于先进机器学习方法和预训练Transformer模型微调的基线模型。实验表明,所提模型在语言识别任务上优于现有最佳预训练模型。相关数据集与代码已开放于https://yashingle-ai.github.io/ILID/ 及Hugging Face开源平台。

原文摘要 · Abstract (English)

The language identification task is a crucial fundamental step in NLP. Often it serves as a pre-processing step for widely used NLP applications such as multilingual machine translation, information retrieval, question and answering, and text summarization. The core challenge of language identification lies in distinguishing languages in noisy, short, and code-mixed environments. This becomes even harder in case of diverse Indian languages that exhibit lexical and phonetic similarities, but have distinct differences. Many Indian languages share the same script, making the task even more challenging. Taking all these challenges into account, we develop and release a dataset of 250K sentences consisting of 23 languages including English and all 22 official Indian languages labeled with their language identifiers, where data in most languages are newly created. We also develop and release baseline models using state-of-the-art approaches in machine learning and fine-tuning pre-trained transformer models. Our models outperforms the state-of-the-art pre-trained transformer models for the language identification task. The dataset and the codes are available at https://yashingle-ai.github.io/ILID/ and in Huggingface open source libraries.

语言识别印度语言多语言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。