用AI模拟中美民众态度变化,找出偏见来源并验证三种去偏方法。
Debiasing International Attitudes: LLM Agents for Simulating US-China Perception Changes
- 构建基于LLM的代理框架,结合新闻与社交媒体数据模拟公众态度
- 魔鬼代言人机制最有效,能显著缓解媒体引发的负面态度
- 发现不同模型存在地域性固有偏见,适合政策与舆情研究者
大型语言模型(LLMs)为计算社会科学中意见演化建模提供了变革性机遇。本研究通过构建基于LLM代理的框架,分析媒体如何影响跨国家态度这一全球极化的关键驱动因素。框架模拟2005至2025年美国公众对中国的态度演变,整合大规模新闻数据与社交媒体画像初始化代理群体,并引入认知感知反思与观点更新机制。提出三种去偏策略:(1) 事实提取,从主观报道中剥离中性事件;(2) 魔鬼代言人代理,模拟批判性语境化;(3) 反事实暴露以揭示模型内在偏见。使用两个前沿LLM(Qwen3-14b和GPT4o)进行仿真,结果显示媒体暴露导致预期的负面态度上升趋势。三种机制均不同程度缓解该趋势,但主观新闻框架仅带来小幅负面影响,而魔鬼代言人代理效果最优,表明中间分析步骤可生成更类人观点。值得注意的是,反事实研究在不同模型间呈现矛盾结果,暗示模型地理来源与其固有偏见相关。研究推进了对基于LLM意见形成与去偏方法的理解,有助于开发更符合人类认知倾向的客观模型。
原文摘要 · Abstract (English)
Large Language Models (LLMs) offer transformative opportunities to address the longstanding challenge of modeling opinion evolution in computational social science. This study investigates how media influences cross-border attitudes - a key driver of global polarization - by developing an LLM-agent framework to disentangle sources of bias and assess LLMs' capacity for human-like opinion formation in response to external information. We introduce an LLM-agent-based framework that models U.S. citizens' attitudes toward China from 2005 to 2025. Our approach integrates large-scale news data with social media profiles to initialize agent populations, which then undergo cognitive-aware reflection and opinion updating. We propose three debiasing mechanisms: (1) fact elicitation, extracting neutral events from subjectively framed news; (2) a devil's advocate agent that simulates critical contextualization; and (3) counterfactual exposure to surface inherent model biases. Simulations with two state-of-the-art LLMs (Qwen3-14b and GPT4o) reveal the expected negative attitudinal trend following media exposure. While all three mechanisms mitigate this trend to varying degrees, results indicate that subjective news framing contributes only modestly to negative attitudes, whereas the devil's advocate agent proves most effective overall, suggesting that intermediate analytical steps can produce more human-like agent opinions. Notably, the counterfactual study reveals contradictory findings across models, suggesting region-specific inherent biases tied to models' geographic origins. By advancing understanding of LLM-based opinion formation and debiasing methods, this study contributes to developing more objective models that better align with human cognitive tendencies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。