研究话题情感如何影响政治立场感知,发现大模型标注会高估效应大小。
Does Topic Sentiment Cause Perceived Ideology? Comparing Human and LLM Annotations in Political News Articles

- 用双重机器学习和中介分析比较人类与大模型标注的意识形态判断
- 零样本大模型普遍夸大情感对立场的影响程度,微调后有所缓解
- 适合关注大模型在因果推断中可靠性的人士参考
我们探讨话题情感是否对感知到的政治意识形态具有因果影响,以及这一结论是否依赖于标注者身份。基于AllSides新闻文章数据,结合Llama-3.3-70b-versatile提供的共享情感标注,比较了专家人类标注者、GPT-4o-mini(基础版与微调版)以及Llama-3.3-70B的意识形态标注结果。采用双重机器学习(DML)和中介分析方法,在四种标注范式下进行评估。结果显示,零样本大模型通常比人类标注高估效应规模,而微调常能将其拉回接近人类水平。该结果对使用大模型标注作为银标签或人类判断代理以开展下游因果分析具有启示意义:其可能可靠反映党派性话题中效应的存在与方向,但无法准确反映效应幅度,可能导致特定话题下意识形态判断的过量或低估。
原文摘要 · Abstract (English)
We ask whether topic sentiment has a causal effect on perceived political ideology, and whether the answer depends on who assigns the ideology label. Using articles from AllSides, paired with shared sentiment annotations from Llama-3.3-70b-versatile, we compare ideology labels from expert human annotators, GPT-4o-mini (baseline and finetuned), and Llama-3.3-70B. We apply Double Machine Learning (DML) and mediation analysis across all four annotation paradigms. Zero-shot LLMs regularly inflate effect sizes relative to human annotations, while fine-tuning often attenuates them back toward the human scale. Our results have implications for the use of LLM annotations as silver labels and as proxies for human judgment in downstream causal analyses: they may be reliable for recovering the presence and direction of effects on the partisan topics, but not their magnitude, leading to over- or under-prediction of some ideology given particular topics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。