arXiv:2604.08566cs.CLcs.LG2026-04

对比大模型与阿拉伯语BERT在战事新闻情感分析中的偏差差异。

Sentiment Classification of Gaza War Headlines: A Comparative Analysis of Large Language Models and Arabic Fine-Tuned BERT Models

  • 用信息论指标量化模型间情感分布差异
  • BERT模型倾向中性,大模型普遍放大负面情绪
  • 适合关注算法偏见与战争叙事的研究者

本研究以2023年加沙战争为案例,分析不同人工智能架构对冲突相关媒体话语的情感解读。基于10,990条阿拉伯语新闻标题(Eleraqi 2026),比较了三个大语言模型与六个微调的阿拉伯语BERT模型。研究不依赖单一人工标注标准,而是将情感分类视为模型架构生成的阐释行为。通过香农熵、Jensen-Shannon距离及偏离集体行为的方差分数等分布性指标,揭示模型间存在显著且非随机的情感分布差异。微调的BERT模型(尤其是MARBERT)强烈偏向中性判断,而大模型则持续放大负面情感,其中LLaMA-3.1-8B几乎完全坍缩为负面。框架条件分析显示,GPT-4.1能根据叙述框架(如人道主义、法律、安全)调整情感判断,其他大模型则表现出有限的情境调节能力。结果表明,模型选择即选择阐释视角,决定冲突叙事的算法建构与情感评估。研究推动媒体研究与计算社会科学的发展,强调算法差异本身应成为分析对象,并警示在战时与危机情境下,不可将自动情感输出视为中立或可互换的媒体基调度量。

原文摘要 · Abstract (English)

This study examines how different artificial intelligence architectures interpret sentiment in conflict-related media discourse, using the 2023 Gaza War as a case study. Drawing on a corpus of 10,990 Arabic news headlines (Eleraqi 2026), the research conducts a comparative analysis between three large language models and six fine-tuned Arabic BERT models. Rather than evaluating accuracy against a single human-annotated gold standard, the study adopts an epistemological approach that treats sentiment classification as an interpretive act produced by model architectures. To quantify systematic differences across models, the analysis employs information-theoretic and distributional metrics, including Shannon Entropy, Jensen-Shannon Distance, and a Variance Score measuring deviation from aggregate model behavior. The results reveal pronounced and non-random divergence in sentiment distributions. Fine-tuned BERT models, particularly MARBERT, exhibit a strong bias toward neutral classifications, while LLMs consistently amplify negative sentiment, with LLaMA-3.1-8B showing near-total collapse into negativity. Frame-conditioned analysis further demonstrates that GPT-4.1 adjusts sentiment judgments in line with narrative frames (e.g., humanitarian, legal, security), whereas other LLMs display limited contextual modulation. These findings suggest that the choice of model constitutes a choice of interpretive lens, shaping how conflict narratives are algorithmically framed and emotionally evaluated. The study contributes to media studies and computational social science by foregrounding algorithmic discrepancy as an object of analysis and by highlighting the risks of treating automated sentiment outputs as neutral or interchangeable measures of media tone in contexts of war and crisis.

情感分析大模型偏见阿拉伯语NLP战争叙事

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。