LLMs能生成以假乱真的多人社交对话,人类辨识率仅略超随机。
The Collective Turing Test: Large Language Models Can Generate Realistic Multi-User Discussions
- 用Llama 3和GPT-4o模拟Reddit真实对话场景。
- 人类误判AI对话为真人内容达39%,对Llama 3的识别准确率仅56%。
- 适合研究社会模拟或评估内容政策影响的学者使用。
大语言模型(LLMs)为模拟在线社区和社交媒体提供了新途径,可用于测试内容推荐算法设计、评估内容政策与干预措施的影响。然而,利用LLMs模拟多用户社交对话的有效性尚未得到充分验证。我们评估了LLMs在社交媒体上模拟真实群体对话的能力。从Reddit收集真实人类对话,并使用Llama 3 70B和GPT-4o在相同主题下生成人工对话。当与真人对话并列呈现给研究参与者时,有39%的案例中,参与者将生成对话误认为是真人创作。特别是针对Llama 3生成的内容,参与者正确识别为AI生成的仅56%,接近随机水平。结果表明,当前LLMs可生成足够逼真的社交媒体对话,足以欺骗人类读者,既展示了社会模拟的应用前景,也警示了其被用于制造虚假社交内容的风险。
原文摘要 · Abstract (English)
Large Language Models (LLMs) offer new avenues to simulate online communities and social media. Potential applications range from testing the design of content recommendation algorithms to estimating the effects of content policies and interventions. However, the validity of using LLMs to simulate conversations between various users remains largely untested. We evaluated whether LLMs can convincingly mimic human group conversations on social media. We collected authentic human conversations from Reddit and generated artificial conversations on the same topic with two LLMs: Llama 3 70B and GPT-4o. When presented side-by-side to study participants, LLM-generated conversations were mistaken for human-created content 39\% of the time. In particular, when evaluating conversations generated by Llama 3, participants correctly identified them as AI-generated only 56\% of the time, barely better than random chance. Our study demonstrates that LLMs can generate social media conversations sufficiently realistic to deceive humans when reading them, highlighting both a promising potential for social simulation and a warning message about the potential misuse of LLMs to generate new inauthentic social media content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。