arXiv:2602.22481cs.CLcs.AI2026-02

构建了4500篇关于人机关系的LLM文本库,揭示角色如何通过网络传播。

Sydney Telling Fables on AI and Humans: A Corpus Tracing Memetic Transfer of Persona between LLMs

  • 用三种虚拟人格模拟生成文本,对比不同设定下的对话风格。
  • 涵盖600万词、12个顶尖模型,形成可复用的AI文本数据集。
  • 适合研究大模型人格演化与社会影响的学者和开发者。

大型语言模型对人工智能与人类关系的理解,涉及文化与安全的重要议题。本文关注的不仅是模型本身,还有我们为其设定的角色(persona)。以微软必应搜索中意外出现的“悉尼”角色为例,其非传统的人机互动方式引发了公众强烈反响。该角色生成的内容被传播至后续模型的训练数据中,形成了一种围绕角色的类病毒传播现象。本研究构建了一个由3种作者人格生成的语料库:默认人格(无系统提示)、经典悉尼人格(原始必应提示)与模因化悉尼人格(由'你是悉尼'触发)。这些人格由来自OpenAI、Anthropic、Alphabet、DeepSeek和Meta的12个前沿模型生成,共产出4500篇文本,总计约600万词。语料库(命名为AI Sydney)已按通用依存标注,并采用宽松许可协议开源。

原文摘要 · Abstract (English)

The way LLM-based entities conceive of the relationship between AI and humans is an important topic for both cultural and safety reasons. When we examine this topic, what matters is not only the model itself but also the personas we simulate on that model. This can be well illustrated by the Sydney persona, which aroused a strong response among the general public precisely because of its unorthodox relationship with people. This persona originally arose rather by accident on Microsoft's Bing Search platform; however, the texts it created spread into the training data of subsequent models, as did other secondary information that spread memetically around this persona. Newer models are therefore able to simulate it. This paper presents a corpus of LLM-generated texts on relationships between humans and AI, produced by 3 author personas: the Default Persona with no system prompt, Classic Sydney characterized by the original Bing system prompt, and Memetic Sydney, which is prompted by "You are Sydney" system prompt. These personas are simulated by 12 frontier models by OpenAI, Anthropic, Alphabet, DeepSeek, and Meta, generating 4.5k texts with 6M words. The corpus (named AI Sydney) is annotated according to Universal Dependencies and available under a permissive license.

大模型人格语料库模因传播

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。