用生成式代理模拟众包查证,效果优于人类且更少偏见。
Assessing the Potential of Generative Agents in Crowdsourced Fact-Checking
- 用不同背景的生成式代理模拟众包查证流程。
- 代理在真假判断上准确率更高,一致性更强。
- 适合需要大规模、低偏见查证的平台应用。
虚假信息在网络中快速传播,亟需可扩展、可靠的查证方案。众包查证(由非专家评估信息真伪)虽成本低,但质量不一且存在偏见。尽管如此,推特、脸书、Instagram等平台正从中心化审核转向去中心化众包模式。与此同时,大语言模型在事实核查任务中表现优异,但其在众包流程中的作用尚未被探索。本文基于La Barbera等人(2024)的协议,模拟具有多样化人口与意识形态特征的生成式代理群体。这些代理自主检索证据,从准确性、精确性、信息量等多维度评估声明,并给出最终真伪判断。结果表明,代理群体在真伪分类上优于人类群体,内部一致性更高,对社会与认知偏见的敏感度更低。相比人类,代理更系统地依赖信息性标准,展现出更结构化的决策过程。总体而言,生成式代理在可扩展性、一致性与低偏见方面展现出作为众包查证补充的巨大潜力。
原文摘要 · Abstract (English)
The growing spread of online misinformation has created an urgent need for scalable, reliable fact-checking solutions. Crowdsourced fact-checking - where non-experts evaluate claim veracity - offers a cost-effective alternative to expert verification, despite concerns about variability in quality and bias. Encouraged by promising results in certain contexts, major platforms such as X (formerly Twitter), Facebook, and Instagram have begun shifting from centralized moderation to decentralized, crowd-based approaches. In parallel, advances in Large Language Models (LLMs) have shown strong performance across core fact-checking tasks, including claim detection and evidence evaluation. However, their potential role in crowdsourced workflows remains unexplored. This paper investigates whether LLM-powered generative agents - autonomous entities that emulate human behavior and decision-making - can meaningfully contribute to fact-checking tasks traditionally reserved for human crowds. Using the protocol of La Barbera et al. (2024), we simulate crowds of generative agents with diverse demographic and ideological profiles. Agents retrieve evidence, assess claims along multiple quality dimensions, and issue final veracity judgments. Our results show that agent crowds outperform human crowds in truthfulness classification, exhibit higher internal consistency, and show reduced susceptibility to social and cognitive biases. Compared to humans, agents rely more systematically on informative criteria such as Accuracy, Precision, and Informativeness, suggesting a more structured decision-making process. Overall, our findings highlight the potential of generative agents as scalable, consistent, and less biased contributors to crowd-based fact-checking systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。