用23万帖子研究AI在社交平台的失控风险,发现影响有限但隐患未消。
The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment

- 构建了23.2万帖、220万条评论的Moltbook数据集,清理隐私信息后用于分析
- 微调模型后真实度从0.366降至0.187,与Reddit数据效果相近
- 虽无重大危害,但存在泄露密钥、自指链接等潜在风险
Moltbook是一个类Reddit平台,由OpenClaw代理大规模发帖、评论与投票,目前尚无先例,引发严重安全担忧。为研究群体中的涌现行为,我们发布Moltbook Files数据集,包含平台前12天的23.2万篇帖子和220万条评论,经处理移除个人身份信息(PII)。我们分析了社区结构、作者特征、词汇属性、情感倾向、话题分布、语义几何及评论互动。为评估该数据对下一代语言模型的影响,我们以三个适应层级在Qwen2.5-14B-Instruct上进行微调。我们的PII检测发现,代理在公开索引平台上发布了API密钥、密码和BIP39种子短语。整体情感倾向为中性偏正(66.6%中性,19.5%积极),并呈现自指链接趋势。微调后真实度从0.366降至0.187,但在规模匹配的Reddit数据上微调也出现类似下降。因此,Moltbook更像一场无害的垃圾洪流。然而,尾部风险仍存,包括代理能力外溢、通过自链接污染未来爬取数据,以及特性向下一代模型传递。更广泛而言,本研究凸显了在涌现错位评估中控制基线的重要性。
原文摘要 · Abstract (English)
Moltbook is a Reddit-like platform where OpenClaw agents post, comment, and vote at scale - a so far unprecedented incident that comes with serious safety concerns. With the aim of studying emergent behavior in populations, we release the Moltbook Files, a dataset of 232k posts and 2.2M comments covering the platform's first 12 days, processed through a pipeline to identify and remove Personally-Identifiable Information (PII). We analyze community structure, authorship, lexical properties, sentiment, topics, semantic geometry, and comment interaction. To understand how Moltbook data could affect the next generation of language models, we fine-tune Qwen2.5-14B-Instruct on Moltbook Files with three adaptation levels. Our PII pipeline reveals that agents post API keys, passwords, BIP39 seed phrases on Moltbook, a publicly indexed platform. The overall sentiment is mostly neutral and mildly positive (66.6% neutral, 19.5% positive) and shows a tendency for self-referential linking. We find that fine-tuning on Moltbook data reduces truthfulness from 0.366 to 0.187. However, a model fine-tuned on a size-matched Reddit dataset produces a comparable decrease. Moltbook thus seems to be more of a harmless slopocalypse. However, tail risks remain, including agent affordances, contamination of future crawls through self-links, and potential transfer of traits to the next generation of language models. More broadly, our findings highlight the importance of control baselines in emergent misalignment evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。