测试大模型模拟网民立场对语境变化的敏感度,发现其立场易受无关信息影响。
Revising Context, Shifting Simulated Stance: Auditing LLM-Based Stance Simulation in Online Discussions

- 通过改写对话上下文,检验大模型对用户立场的模拟稳定性。
- 文本与图文混合改写均引发显著立场偏移,平均方向变动超30%。
- 适合关注大模型社会模拟风险的研究者与政策制定者参考。
大型语言模型被广泛用于模拟社交媒体用户并推断个体在在线讨论中的反应。然而,这些模拟是否准确反映特定用户的信念,或是否对语义上无关的对话上下文变化高度敏感,仍不明确。本文提出反事实上下文改写框架,用于审计基于LLM的立场模拟效果。给定原始在线对话,先推断目标用户对某话题的立场,再通过受控策略修改对话上下文,重新模拟用户立场。对比纯文本与融合表情包的多模态改写方式,评估平均立场方向偏移和立场转变率两个核心指标。结果显示,在不同极化偏好机制下,文本与多模态策略均能引发有效且稳定的立场转变。本研究贡献了一个评估LLM立场模拟上下文敏感性的框架(https://github.com/STARResearchLab/StanceShift),揭示了使用大模型作为社交媒体用户代理进行社会模拟的潜力与风险。
原文摘要 · Abstract (English)
Large language models are increasingly used to simulate social media users and infer how individuals may respond to online discussions. However, it remains unclear whether these simulations reflect precise user-specific beliefs or whether they are highly sensitive to semantically independent changes in conversational contexts. In this work, we study counterfactual context revision as a framework for auditing LLM-based stance simulation. Given an original online conversation, we first infer a target user's stance toward a specific topic. We then apply controlled revision strategies to the conversational context and simulate the user's stance again under the revised context. We compare text-only revision strategies with a multimodal one that incorporates meme-based context and evaluate two main effectiveness metrics, i.e., average directional stance shift and stance transition rate. The results reveal effective and robust stance transitions in both text-only and multimodal strategies across different polarization-preference mechanisms. Our study contributes an evaluation framework for understanding the context sensitivity of LLM-based stance simulation (https://github.com/STARResearchLab/StanceShift). More broadly, it highlights both the promise and risk of using LLMs as proxies for social media users in social simulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。