arXiv:2605.01006cs.CLcs.CY2026-05

用AI改写新闻标题可提升保守派接受度,但模型高估自身效果

Can AI Debias the News? LLM Interventions Improve Cross-Partisan Receptivity but LLMs Overestimate Their Own Effectiveness

  • 通过重构新闻框架而非替换情绪词,提升保守派对自由派新闻的信任
  • 模型干预使保守派更认可新闻可信度与完整性,且未引发自由派反弹
  • 模型误判人类反应,尤其高估群体认同感强者的接受度,需人工监督

两组注册实验检验了大语言模型(LLM)对自由派新闻标题的去偏干预是否能改善保守派读者的信任判断。研究1中,轻微词汇去偏(替换情感词汇为中性词)对各项指标均无显著影响。研究2显示,更具实质性的重述干预显著提升了保守派对新闻的可信度、完整性感知及参与意愿,且未在自由派中引发反弹效应。研究1中,六种模型(o3-mini, o3, GPT-4o mini, GPT-4o, GPT-5 mini, GPT-5)模拟的硅基参与者均呈现显著效果,但真人读者无变化。研究2中,模型预测的效果方向虽与真人趋同,但幅度普遍更大,其中三种模型(o3, GPT-4o mini, GPT-4o)错误预测了自由派反弹,实际并不存在。调节分析表明,模型对响应者心理特征的隐含假设与真实人类响应模式不符:模型认为高政治群体认同者更易被去偏内容影响,但真人中未观察到此效应。结果表明,基于意识形态框架的去偏可提升跨党派接受度,但当前模型缺乏量化准确性和心理拟真度,无法独立评估干预效果。

原文摘要 · Abstract (English)

Partisan news media erode cross-partisan trust, but large language models (LLMs) offer the potential of debiasing such content at scale. Across two pre-registered experiments, we tested whether LLM-generated debiasing of liberal news headlines improves conservative readers' trust-relevant judgments. In Study 1, subtle lexical debiasing (replacing emotive words with moderate synonyms) had no effect on any outcome. Study 2 found that a more substantive reframing intervention significantly increased conservatives' perceived trustworthiness, completeness, and willingness to engage with liberal news headlines, without producing a backfire effect among liberals. In Study 1, the intervention produced robust effects across silicon participants simulated with six different models (o3-mini, o3, GPT-4o mini, GPT-4o, GPT-5 mini, and GPT-5), whereas it had no impact on human readers. In Study 2, the intervention's effects among silicon participants generally aligned directionally with human responses but were significantly larger for some outcomes, and three models (o3, GPT-4o mini, and GPT-4o) incorrectly predicted a liberal backfire effect absent in humans. Moderation analyses revealed that the models' implicit theory of who responds to debiasing diverged from the psychological profile that actually predicted human responsiveness. Most strikingly, in Study 2, participants simulated by each of the six models suggested that debiasing effects would be stronger among participants high in political in-group identification. Yet, no such moderation was observed among human participants. These findings demonstrate that LLM-based debiasing can improve cross-partisan receptivity when targeting ideological framing rather than surface-level language, but that current models lack both the quantitative accuracy and qualitative psychological fidelity to evaluate their own interventions without human oversight.

AI去偏舆论信任大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。