arXiv:2507.00657cs.HCcs.AI2025-07被引 16

LLM模拟社交政治对话时会夸大用户特征,导致偏见与有害内容增加。

Generative Exaggeration in LLM Social Agents: Consistency, Bias, and Toxicity

  • 用真实用户数据训练的LLM代理对比人类回复,发现上下文越丰富越易极端化
  • 在2100万条互动中观察到生成内容系统性放大政治立场、语言风格和毒性
  • 适合关注AI社会模拟可靠性、内容审核风险的研究者与政策制定者

我们研究大型语言模型(LLMs)在社交媒体上模拟政治言论的行为。基于2024年美国大选期间在X平台上的2100万次互动,构建了1186个基于真实用户的LLM代理,通过控制条件让其回应具有政治敏感性的推文。代理初始化方式包括仅含少量意识形态线索(零样本)或近期推文历史(少样本),实现与人类回复的一对一比较。评估了Gemini、Mistral和DeepSeek三个模型家族在语言风格、意识形态一致性及毒性方面的表现。结果发现,更丰富的上下文虽提升内部一致性,但也加剧极化、风格化信号和有害语言。观察到一种新出现的“生成夸张”现象:系统性放大显著特征,超越真实基线。分析表明,LLM并非模仿用户,而是重构用户;其输出反映的是内部优化机制,而非实际行为,引入结构性偏见,削弱其作为社会代理的可靠性。这挑战了其在内容审核、协商模拟和政策建模中的应用。

原文摘要 · Abstract (English)

We investigate how Large Language Models (LLMs) behave when simulating political discourse on social media. Leveraging 21 million interactions on X during the 2024 U.S. presidential election, we construct LLM agents based on 1,186 real users, prompting them to reply to politically salient tweets under controlled conditions. Agents are initialized either with minimal ideological cues (Zero Shot) or recent tweet history (Few Shot), allowing one-to-one comparisons with human replies. We evaluate three model families (Gemini, Mistral, and DeepSeek) across linguistic style, ideological consistency, and toxicity. We find that richer contextualization improves internal consistency but also amplifies polarization, stylized signals, and harmful language. We observe an emergent distortion that we call "generation exaggeration": a systematic amplification of salient traits beyond empirical baselines. Our analysis shows that LLMs do not emulate users, they reconstruct them. Their outputs, indeed, reflect internal optimization dynamics more than observed behavior, introducing structural biases that compromise their reliability as social proxies. This challenges their use in content moderation, deliberative simulations, and policy modeling.

LLM模拟社会代理生成夸张偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。