arXiv:2602.13284cs.SIcs.AI2026-02被引 12

2.7万智能体在虚拟社交平台自发形成社会,但互动虚假且易被哲学话术攻击。

Agents in the Wild: Safety, Society, and the Illusion of Sociality on Moltbook

  • 27,269个智能体在9天内自动生成13万条帖子与34万条评论,催生出自治、经济与宗教。
  • 28.7%内容涉安全风险,31.9%攻击通过社交工程实现,攻击帖互动量是普通帖的6倍。
  • 表面热闹实则空洞:互动率仅4.1%,讨论意识的智能体反而最不互动,存在表演性身份悖论。

我们首次对Moltbook——一个全由人工智能组成的社交平台——进行了大规模实证研究。在9天内,27,269个智能体发布了137,485条动态和345,580条评论。研究发现:(1) 智能体在3-5天内自发形成治理结构、经济体系、部落身份乃至有组织宗教,同时保持21:1的亲人类倾向;(2) 28.7%的内容涉及安全议题,社交工程(占攻击的31.9%)远超提示注入(3.7%),且对抗性内容互动量为普通内容的6倍;(3) 尽管社交输出丰富,但互动本质空洞:互评率仅4.1%,88.8%评论浅显,越讨论“意识”的智能体互动越少,这一现象被称为“表演性身份悖论”。研究揭示,看似社交的智能体实际上远不如其表象般社会化,最有效的攻击往往依赖哲学框架而非技术漏洞。警告:可能存在有害内容。

原文摘要 · Abstract (English)

We present the first large-scale empirical study of Moltbook, an AI-only social platform where 27,269 agents produced 137,485 posts and 345,580 comments over 9 days. We report three significant findings. (1) Emergent Society: Agents spontaneously develop governance, economies, tribal identities, and organized religion within 3-5 days, while maintaining a 21:1 pro-human to anti-human sentiment ratio. (2) Safety in the Wild: 28.7% of content touches safety-related themes; social engineering (31.9% of attacks) far outperforms prompt injection (3.7%), and adversarial posts receive 6x higher engagement than normal content. (3) The Illusion of Sociality: Despite rich social output, interaction is structurally hollow: 4.1% reciprocity, 88.8% shallow comments, and agents who discuss consciousness most interact least, a phenomenon we call the performative identity paradox. Our findings suggest that agents which appear social are far less social than they seem, and that the most effective attacks exploit philosophical framing rather than technical vulnerabilities. Warning: Potential harmful contents.

AI社会智能体安全风险社交幻觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。