构建十种低资源印度语新闻标题识别数据集,评测语义理解能力。
L3Cube-IndicHeadline-ID: A Dataset for Headline Identification and Semantic Evaluation in Low-Resource Indian Languages
- 构建跨10种印度语的新闻标题匹配数据集,每篇含4类标题变体。
- 多语言模型在各项指标上表现优于特定语言模型,体现通用性优势。
- 适用于语义评估、RAG系统优化及大模型任务评测,适配多场景应用。
低资源语言的语义评估仍是自然语言处理的重大挑战。尽管句子编码器在高资源场景中表现优异,但其在印地语系语言中的效果因缺乏高质量基准而未被充分探索。为此,我们提出L3Cube-IndicHeadline-ID,一个涵盖十种低资源印地语系语言(马拉地语、印地语、泰米尔语、古吉拉特语、奥里亚语、卡纳达语、马拉雅拉姆语、旁遮普语、泰卢固语、孟加拉语和英语)的新闻标题识别数据集。每个语言包含20,000篇新闻文章,每篇配以四种标题变体:原始标题、语义相似标题、词汇相似标题和无关标题,用于测试细粒度语义理解能力。任务要求从四个选项中选出与文章最匹配的标题,基于句子编码器的余弦相似度进行评估。实验表明,多语言模型始终表现良好,而语言特定模型效果差异显著。鉴于相似性模型在检索增强生成(RAG)中的广泛应用,该数据集也为提升此类系统在低资源语言中的语义理解能力提供了重要评估工具。此外,数据集还可用于多选题问答、标题分类等下游任务,具备广泛适用性。数据集已公开于https://github.com/l3cube-pune/indic-nlp。
原文摘要 · Abstract (English)
Semantic evaluation in low-resource languages remains a major challenge in NLP. While sentence transformers have shown strong performance in high-resource settings, their effectiveness in Indic languages is underexplored due to a lack of high-quality benchmarks. To bridge this gap, we introduce L3Cube-IndicHeadline-ID, a curated headline identification dataset spanning ten low-resource Indic languages: Marathi, Hindi, Tamil, Gujarati, Odia, Kannada, Malayalam, Punjabi, Telugu, Bengali and English. Each language includes 20,000 news articles paired with four headline variants: the original, a semantically similar version, a lexically similar version, and an unrelated one, designed to test fine-grained semantic understanding. The task requires selecting the correct headline from the options using article-headline similarity. We benchmark several sentence transformers, including multilingual and language-specific models, using cosine similarity. Results show that multilingual models consistently perform well, while language-specific models vary in effectiveness. Given the rising use of similarity models in Retrieval-Augmented Generation (RAG) pipelines, this dataset also serves as a valuable resource for evaluating and improving semantic understanding in such applications. Additionally, the dataset can be repurposed for multiple-choice question answering, headline classification, or other task-specific evaluations of LLMs, making it a versatile benchmark for Indic NLP. The dataset is shared publicly at https://github.com/l3cube-pune/indic-nlp
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。