arXiv:2601.06477cs.CLcs.CY2026-01被引 1

构建印度区域偏见数据集,评估大模型识别地区歧视能力

IndRegBias: A Dataset for Studying Indian Regional Biases in English and Code-Mixed Social Media Comments

  • 收集2.5万条印地语与混合语社交媒体评论,标注区域偏见程度
  • 微调显著提升大模型对印度区域偏见及严重性的识别准确率
  • 为研究印度地区歧视提供可复用数据集,适合偏见检测研究者

本文构建了名为IndRegBias的数据集,聚焦印度社交媒体评论中的区域偏见问题。研究选取25,000条来自Reddit和YouTube的热门话题评论,涵盖印度各地的区域性议题。提出多层级标注策略,评估评论中区域偏见的严重程度。通过零样本、少样本及微调方法,测试开源大语言模型(LLMs)和印地语语言模型(ILMs)对区域偏见的检测能力。结果表明,零样本与少样本方法在多数模型上表现不佳;而微调显著提升了模型对印度区域偏见及其严重性的识别性能。

原文摘要 · Abstract (English)

Warning: This paper consists of examples representing regional biases in Indian regions that might be offensive towards a particular region. While social biases corresponding to gender, race, socio-economic conditions, etc., have been extensively studied in the major applications of Natural Language Processing (NLP), biases corresponding to regions have garnered less attention. This is mainly because of (i) difficulty in the extraction of regional bias datasets, (ii) disagreements in annotation due to inherent human biases, and (iii) regional biases being studied in combination with other types of social biases and often being under-represented. This paper focuses on creating a dataset IndRegBias, consisting of regional biases in an Indian context reflected in users' comments on popular social media platforms, namely Reddit and YouTube. We carefully selected 25,000 comments appearing on various threads in Reddit and videos on YouTube discussing trending topics on regional issues in India. Furthermore, we propose a multilevel annotation strategy to annotate the comments describing the severity of regional biased statements. To detect the presence of regional bias and its severity in IndRegBias, we evaluate open-source Large Language Models (LLMs) and Indic Language Models (ILMs) using zero-shot, few-shot, and fine-tuning strategies. We observe that zero-shot and few-shot approaches show lower accuracy in detecting regional biases and severity in the majority of the LLMs and ILMs. However, the fine-tuning approach significantly enhances the performance of the LLM in detecting Indian regional bias along with its severity.

区域偏见数据集大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。