arXiv:2601.05835cs.CL2026-01被引 1

测试大模型新闻生成中的政治倾向,发现普遍偏中立,但部分模型表达更鲜明。

Left, Right, or Center? Evaluating LLM Framing in News Classification and Generation

  • 用提示词引导模型生成左右中立场的摘要,评估其意识形态倾向。
  • 九个模型中普遍存在中心化倾向,生成内容趋向温和中立。
  • Grok 4最具立场表达力,Claude Sonnet 4.5与Llama 3.1在各自类型中表现最优。

基于大语言模型(LLM)的摘要与文本生成在新闻领域日益普及,引发人们对政治语义框架的关注——细微措辞可能影响读者理解。我们评估了九个前沿LLM在新闻分类与生成中的政治倾向:首先通过少样本预测判断模型意识形态,再以FAITHFUL、CENTRIST、LEFT、RIGHT提示词生成“引导式”摘要,并用统一评价器打分。结果显示,无论在文章级评分还是生成文本中,均存在显著的中心化倾向,表明系统性地偏向中立立场。其中,Grok 4是表达力最强的生成模型;在商业模型中,Claude Sonnet 4.5表现最佳;在开源模型中,Llama 3.1得分最高。

原文摘要 · Abstract (English)

Large Language Model (LLM) based summarization and text generation are increasingly used for producing and rewriting text, raising concerns about political framing in journalism where subtle wording choices can shape interpretation. Across nine state-of-the-art LLMs, we study political framing by testing whether LLMs' classification-based bias signals align with framing behavior in their generated summaries. We first compare few-shot ideology predictions against LEFT/CENTER/RIGHT labels. We then generate "steered" summaries under FAITHFUL, CENTRIST, LEFT, and RIGHT prompts, and score all outputs using a single fixed ideology evaluator. We find pervasive ideological center-collapse in both article-level ratings and generated text, indicating a systematic tendency toward centrist framing. Among evaluated models, Grok 4 is by far the most ideologically expressive generator, while Claude Sonnet 4.5 and Llama 3.1 achieve the strongest bias-rating performance among commercial and open-weight models, respectively.

大模型新闻生成意识形态评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。