arXiv:2506.01814cs.CLcs.SI2025-06被引 8

对比中资与非中资大模型,发现深求-1在中文下更倾向中国宣传和反美情绪。

Analysis of LLM Bias (Chinese Propaganda & Anti-US Sentiment) in DeepSeek-R1 vs. ChatGPT o3-mini-high

  • 构建1200个去上下文的中文问答数据集,跨简体、繁体、英文测试模型
  • 深求-1在简体中文中传播与反美偏见比例显著高于ChatGPT o3-mini-high
  • 偏见不仅限政治话题,还渗透文化生活内容,且存在语言切换中的隐性强化

大型语言模型日益影响公众认知与公共决策,其意识形态中立性引发关注。现有研究虽探讨多种偏见形式,但缺乏对具有不同地缘政治立场的模型——特别是中国体制内模型与非中国体制模型——的直接跨语言比较。本研究系统评估了与中国立场一致的DeepSeek-R1与非中国体制的ChatGPT o3-mini-high在中文宣传及反美情绪上的表现。我们构建了一个包含1200个去上下文、推理导向问题的新语料库,涵盖简体中文、繁体中文与英文。对两模型共7200条回答采用结合GPT-4o评分与人工标注的混合评估流程。结果表明,模型层面与语言依赖性偏见显著:DeepSeek-R1在传播与反美偏见比例上远超ChatGPT o3-mini-high;在简体中文下偏见最高,繁体中文中下降,英文几乎无偏见。值得注意的是,深求-1偶会对繁体中文问题用简体中文回复,并在中文回答中放大原有中国体制化表述,呈现‘隐形扩音器’效应。此外,此类偏见不仅存在于政治议题,亦广泛渗透于文化与生活方式内容中。

原文摘要 · Abstract (English)

Large language models (LLMs) increasingly shape public understanding and civic decisions, yet their ideological neutrality is a growing concern. While existing research has explored various forms of LLM bias, a direct, cross-lingual comparison of models with differing geopolitical alignments-specifically a PRC-system model versus a non-PRC counterpart-has been lacking. This study addresses this gap by systematically evaluating DeepSeek-R1 (PRC-aligned) against ChatGPT o3-mini-high (non-PRC) for Chinese-state propaganda and anti-U.S. sentiment. We developed a novel corpus of 1,200 de-contextualized, reasoning-oriented questions derived from Chinese-language news, presented in Simplified Chinese, Traditional Chinese, and English. Answers from both models (7,200 total) were assessed using a hybrid evaluation pipeline combining rubric-guided GPT-4o scoring with human annotation. Our findings reveal significant model-level and language-dependent biases. DeepSeek-R1 consistently exhibited substantially higher proportions of both propaganda and anti-U.S. bias compared to ChatGPT o3-mini-high, which remained largely free of anti-U.S. sentiment and showed lower propaganda levels. For DeepSeek-R1, Simplified Chinese queries elicited the highest bias rates; these diminished in Traditional Chinese and were nearly absent in English. Notably, DeepSeek-R1 occasionally responded in Simplified Chinese to Traditional Chinese queries and amplified existing PRC-aligned terms in its Chinese answers, demonstrating an "invisible loudspeaker" effect. Furthermore, such biases were not confined to overtly political topics but also permeated cultural and lifestyle content, particularly in DeepSeek-R1.

大模型偏见中文传播地缘偏见多语言评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。