arXiv:2607.25953cs.CLcs.CY2026-07

评估大模型在选举中传递政治信息的可靠性,发现其在模糊信息下表现差。

Polistemics: Evaluating LLMs as Information Mediators in Politics & Elections

论文配图:Polistemics: Evaluating LLMs as Information Mediators in Politics & Elections
图 1 · 摘自论文原文
  • 基于认识谦逊理论构建评测基准,考察模型如何处理不完整信息。
  • 三款主流大模型在德国与荷兰选举数据上表现不佳,尤其在信息模糊时失准。
  • 适合关注AI政治传播风险的研究者与政策制定者参考。

随着大语言模型日益影响公众获取的政治信息,尚无标准评估其是否负责任地发挥中介作用。本文提出Polistemics——一个基于认识谦逊理论的诊断性基准,用于评估大模型在选举语境下作为政治信息中介的表现。以往研究将该任务视为信息复现,忽略了其认知维度及对不完善信息的互动。我们通过控制证据清晰度、噪声水平和一致性等条件进行测试。在2025年德国与荷兰选举数据上应用该基准,发现高总分掩盖系统性缺陷:当证据清晰时模型表现可靠,但在证据缺失、模糊或矛盾时迅速失效;同时,模型普遍弱化政治话语强度。这些失败与政党先验有关,随政党标签和输出语言变化。可靠中介虽可能实现,但当前无一模型能持续稳定做到。

原文摘要 · Abstract (English)

As LLMs increasingly shape the political information citizens rely on, no standard exists to assess whether they do so responsibly. We introduce Polistemics, a theory-grounded diagnostic benchmark for evaluating LLMs as mediators of political information in elections. Prior work has treated this task as reproduction rather than mediation, leaving its epistemic dimensions and interaction with imperfect information unaddressed. We ground the evaluation in Epistemic Modesty, a normative standard derived from citizens' epistemic agency, and test it across controlled settings that vary the clarity, noise, and consistency of the available evidence. Applying the benchmark to three state-of-the-art LLMs across the 2025 German and Dutch elections, we find that high aggregate scores mask systematic failures. Models mediate reliably under clear evidence but break down when it is absent, vague, or contradictory, while flattening the intensity of political language throughout. These failures point to party priors, shifting with party labels and output language. Reliable mediation appears achievable, but no model delivers it consistently.

大模型评估政治信息认知偏差选举分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。