arXiv:2510.03898cs.CL2025-10中稿 · BLP at AACL-IJCNLP…被引 4

首个孟加拉语新闻政治偏见基准数据集,助力检测媒体立场

Read Between the Lines: A Benchmark for Uncovering Political Bias in Bangla News Articles

  • 构建200篇高争议孟加拉语新闻的标注数据集,区分政府倾向、批评与中立立场
  • 大模型对政府批评内容识别准确率最高达0.83,但对中立文章几乎失效(最低F1 0.00)
  • 揭示模型在模糊叙事中易误判为政府倾向,适合低资源语言研究者使用

检测媒体偏见在南亚地区至关重要,但针对孟加拉语的政治偏见标注数据集和计算研究仍极为稀缺。由于孟加拉语政治立场识别需理解语言线索、文化背景、隐含情绪、修辞策略及语码转换等复杂因素,我们提出首个基准数据集,包含200篇具有重大政治意义且高度争议的孟加拉语新闻文章,分别标注为政府倾向、政府批判与中立立场,并提供用于评估大语言模型(LLMs)的诊断分析工具。对28个专有及开源大模型的全面评估显示,其在检测政府批判内容时表现良好(最高F1达0.83),但在识别中立文章时表现极差(最低F1为0.00),且普遍倾向于过度预测政府倾向,常将模糊叙述误判为支持立场。该数据集及其诊断工具为推进孟加拉语媒体立场检测研究奠定了基础,并为提升低资源语言中大模型性能提供了关键洞察。

原文摘要 · Abstract (English)

Detecting media bias is crucial, specifically in the South Asian region. Despite this, annotated datasets and computational studies for Bangla political bias research remain scarce. Crucially because, political stance detection in Bangla news requires understanding of linguistic cues, cultural context, subtle biases, rhetorical strategies, code-switching, implicit sentiment, and socio-political background. To address this, we introduce the first benchmark dataset of 200 politically significant and highly debated Bangla news articles, labeled for government-leaning, government-critique, and neutral stances, alongside diagnostic analyses for evaluating large language models (LLMs). Our comprehensive evaluation of 28 proprietary and open-source LLMs shows strong performance in detecting government-critique content (F1 up to 0.83) but substantial difficulty with neutral articles (F1 as low as 0.00). Models also tend to over-predict government-leaning stances, often misinterpreting ambiguous narratives. This dataset and its associated diagnostics provide a foundation for advancing stance detection in Bangla media research and offer insights for improving LLM performance in low-resource languages.

政治偏见孟加拉语大模型评估低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。