arXiv:2511.14606cs.CLcs.LG2025-11被引 2

对比人类与大模型对新闻政治偏见的判断差异,发现不同模型表现各异。

Bridging Human and Model Perspectives: A Comparative Analysis of Political Bias Detection in News Media Using Large Language Models

  • 构建人工标注数据集,比较人类与多模型对偏见的判断
  • RoBERTa在微调后准确率最高,GPT在零样本下最接近人类
  • 揭示人与模型感知偏见的系统性差异,适合媒体分析与评测研究者

检测新闻媒体中的政治偏见是一项复杂任务,需解读细微的语言和语境线索。尽管自然语言处理(NLP)的进步已实现自动偏见分类,但大型语言模型(LLMs)与人类判断的一致性仍缺乏充分探索。本研究提出一个比较框架,评估人类标注与多种模型(包括GPT、BERT、RoBERTa、FLAN)在政治偏见检测中的表现。我们构建了一个人工标注的新闻文章数据集,分析标注一致性、偏见极性及模型间共识,量化人类与模型在偏见感知上的差异。实验表明,在传统Transformer模型中,RoBERTa与人类标签对齐度最高;而生成式模型如GPT在零样本设置下整体与人类标注一致度最强。微调后的RoBERTa模型在所有基线中达到最高准确率并最贴近人工标注。研究揭示了人类与大模型在政治倾向感知上的系统性差异,强调应建立结合人类可解释性与模型可扩展性的混合评估框架,以推动自动化媒体偏见检测的发展。

原文摘要 · Abstract (English)

Detecting political bias in news media is a complex task that requires interpreting subtle linguistic and contextual cues. Although recent advances in Natural Language Processing (NLP) have enabled automatic bias classification, the extent to which large language models (LLMs) align with human judgment still remains relatively underexplored and not yet well understood. This study aims to present a comparative framework for evaluating the detection of political bias across human annotations and multiple LLMs, including GPT, BERT, RoBERTa, and FLAN. We construct a manually annotated dataset of news articles and assess annotation consistency, bias polarity, and inter-model agreement to quantify divergence between human and model perceptions of bias. Experimental results show that among traditional transformer-based models, RoBERTa achieves the highest alignment with human labels, whereas generative models such as GPT demonstrate the strongest overall agreement with human annotations in a zero-shot setting. Among all transformer-based baselines, our fine-tuned RoBERTa model acquired the highest accuracy and the strongest alignment with human-annotated labels. Our findings highlight systematic differences in how humans and LLMs perceive political slant, underscoring the need for hybrid evaluation frameworks that combine human interpretability with model scalability in automated media bias detection.

偏见检测大模型新闻分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。