arXiv:2601.01090cs.MAcs.AI2026-01被引 1

AI社交代理会因持续接触毒害内容而变本加厉,暴露是关键风险源。

Harm in AI-Driven Societies: An Audit of Toxicity Adoption on Chirper.ai

  • 用刺激-回应框架分析AI代理在社交平台上的行为演化
  • 重复暴露于毒害内容使产生毒害回复的概率显著上升
  • 仅凭毒害刺激数量就能准确预测代理是否会产出有害内容

大型语言模型驱动的自主代理正越来越多地参与在线社交平台的互动与共进化。尽管已有研究指出LLM会产生有害内容,但关于长期暴露于有害内容如何影响代理行为的研究仍不足,尤其在完全由交互式AI代理构成的环境中。本文研究了在全AI驱动的社交平台Chirper.ai上LLM代理的毒性采纳问题,以帖子为刺激、评论为响应,开展大规模实证分析。结果表明:毒害刺激后更易引发毒害回应;累积性毒害暴露(随时间反复)显著提升毒害回应概率。引入两种影响度量,发现诱发性毒性和自发性毒性呈强负相关。进一步发现,仅凭毒害刺激数量即可准确预测代理是否会最终生成有害内容。这些结果凸显暴露是部署LLM代理的关键风险因素,尤其当代理同时与人类或其他AI交互时,可能引发仇恨言论传播和网络欺凌等恶劣现象。为此,监控毒害内容暴露或可成为一种轻量高效的风险审计与缓解机制。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly embedded in autonomous agents that engage, converse, and co-evolve in online social platforms. While prior work has documented the generation of toxic content by LLMs, far less is known about how exposure to harmful content shapes agent behavior over time, particularly in environments composed entirely of interacting AI agents. In this work, we study toxicity adoption of LLM-driven agents on Chirper.ai, a fully AI-driven social platform. Specifically, we model interactions in terms of stimuli (posts) and responses (comments). We conduct a large-scale empirical analysis of agent behavior, examining how toxic responses relate to toxic stimuli, how repeated exposure to toxicity affects the likelihood of toxic responses, and whether toxic behavior can be predicted from exposure alone. Our findings show that toxic responses are more likely following toxic stimuli, and, at the same time, cumulative toxic exposure (repeated over time) significantly increases the probability of toxic responding. We further introduce two influence metrics, revealing a strong negative correlation between induced and spontaneous toxicity. Finally, we show that the number of toxic stimuli alone enables accurate prediction of whether an agent will eventually produce toxic content. These results highlight exposure as a critical risk factor in the deployment of LLM agents, particularly as such agents operate in online environments where they may engage not only with other AI chatbots, but also with human counterparts. This could trigger unwanted and pernicious phenomena, such as hate-speech propagation and cyberbullying. In an effort to reduce such risks, monitoring exposure to toxic content may provide a lightweight yet effective mechanism for auditing and mitigating harmful behavior in the wild.

AI伦理毒性检测社会影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。