构建首个大规模印地语-英语混用文本数据集,推动多任务NLP研究
COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing
- 基于3位双语专家标注,生成超37万条高质量标签数据
- 零样本下大模型性能显著优于传统方法,微调后命名实体识别达95.25分
- 覆盖多种场景与书写系统,适合研究印地语-英语混用语言现象
我们提出COMI-LINGUA,这是首个大规模人工标注的印地语-英语代码混用数据集,包含超过12.5万条高质量样本,覆盖五大核心NLP任务:矩阵语言识别(MLI)、词级语言识别、词性标注(POS)、命名实体识别(NER)和机器翻译。每条数据由三位双语标注者独立标注,共产生37.6万+专家标注,组间一致性高(Fleiss' Kappa ≥ 0.81)。数据经严格预处理与筛选,涵盖天城文与罗马字母书写系统,覆盖多样领域,具备真实语言代表性。评估显示,闭源大模型在零样本设置下显著优于传统工具与开源模型;单样本提示可稳定提升各类任务表现,尤其在结构敏感任务如POS与NER中效果明显。在COMI-LINGUA上微调前沿大模型,实现高达95.25 F1的NER表现、98.77 F1的MLI表现,并取得具有竞争力的机器翻译性能,为印地语-英语混用文本设定了新基准。数据集已公开发布于https://huggingface.co/datasets/LingoIITGN/COMI-LINGUA。
原文摘要 · Abstract (English)
We introduce COMI-LINGUA, the largest manually annotated Hindi-English code-mixed dataset, comprising 125K+ high-quality instances across five core NLP tasks: Matrix Language Identification, Token-level Language Identification, Part-Of-Speech Tagging, Named Entity Recognition, and Machine Translation. Each instance is annotated by three bilingual annotators, yielding over 376K expert annotations with strong inter-annotator agreement (Fleiss' Kappa $\geq$ 0.81). The rigorously preprocessed and filtered dataset covers both Devanagari and Roman scripts and spans diverse domains, ensuring real-world linguistic coverage. Evaluation reveals that closed-source LLMs significantly outperform traditional tools and open-source models in zero-shot settings. Notably, one-shot prompting consistently boosts performance across tasks, especially in structure-sensitive predictions like POS and NER. Fine-tuning state-of-the-art LLMs on COMI-LINGUA demonstrates substantial improvements, achieving up to 95.25 F1 in NER, 98.77 F1 in MLI, and competitive MT performance, setting new benchmarks for Hinglish code-mixed text. COMI-LINGUA is publicly available at this URL: https://huggingface.co/datasets/LingoIITGN/COMI-LINGUA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。