arXiv:2508.19492cs.CYcs.CL2025-08

对比中西大模型新闻评估差异,发现模型来源影响内容判断结果。

Geopolitical Parallax: Beyond Walter Lippmann Just After Large Language Models

  • 用嵌入向量比较中西模型对新闻的评分差异
  • 西方模型更倾向给巴勒斯坦报道打高主观与正向情绪分
  • 中文模型更强调新颖性与描述性,适合跨文化媒体分析

新闻客观性长期处于中立事实报道理想与主观框架不可避免之间的张力中。随着大语言模型(LLMs)的出现,这种张力由算法系统中介,其训练数据与设计选择可能嵌入文化或意识形态偏见。本研究通过对比中国源(Qwen、BGE、Jina)与西方源(Snowflake、Granite)模型家族的文本嵌入,考察了地缘政治视角偏差——即新闻质量与主观性评估的系统性差异。我们在一个涵盖十五个维度(风格、信息、情感等)的人工标注新闻质量基准上进行评估,并在涉及敏感政治议题的平行语料库(包括巴勒斯坦及中美互评报道)上展开分析。通过逻辑回归探测与匹配主题评估,量化了不同模型族在各类指标上的正类预测概率差异。结果显示,模型来源引发的分歧具有一致性和非随机性:在巴勒斯坦相关报道中,西方模型赋予更高主观性与积极情绪分,而中文模型则更突出新颖性与描述性;跨话题分析显示,中文模型对美国报道的流畅性、简洁性、技术性与整体质量评分显著更低,但负向情绪分更高。这些模式符合媒体偏见理论,并区分出语义、情感与关系层面的主观性,拓展了大模型偏见研究,表明地缘政治框架效应持续存在于下游质量评估任务中。结论指出,基于大模型的媒体评估流程需进行文化校准,以避免将内容差异误判为模型偏见。

原文摘要 · Abstract (English)

Objectivity in journalism has long been contested, oscillating between ideals of neutral, fact-based reporting and the inevitability of subjective framing. With the advent of large language models (LLMs), these tensions are now mediated by algorithmic systems whose training data and design choices may themselves embed cultural or ideological biases. This study investigates geopolitical parallax-systematic divergence in news quality and subjectivity assessments-by comparing article-level embeddings from Chinese-origin (Qwen, BGE, Jina) and Western-origin (Snowflake, Granite) model families. We evaluate both on a human-annotated news quality benchmark spanning fifteen stylistic, informational, and affective dimensions, and on parallel corpora covering politically sensitive topics, including Palestine and reciprocal China-United States coverage. Using logistic regression probes and matched-topic evaluation, we quantify per-metric differences in predicted positive-class probabilities between model families. Our findings reveal consistent, non-random divergences aligned with model origin. In Palestine-related coverage, Western models assign higher subjectivity and positive emotion scores, while Chinese models emphasize novelty and descriptiveness. Cross-topic analysis shows asymmetries in structural quality metrics Chinese-on-US scoring notably lower in fluency, conciseness, technicality, and overall quality-contrasted by higher negative emotion scores. These patterns align with media bias theory and our distinction between semantic, emotional, and relational subjectivity, and extend LLM bias literature by showing that geopolitical framing effects persist in downstream quality assessment tasks. We conclude that LLM-based media evaluation pipelines require cultural calibration to avoid conflating content differences with model-induced bias.

大模型偏见检测地缘政治新闻评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。