用角色扮演游戏测试大模型如何呈现集体记忆
ROBOPSY PL[AI]: Using Role-Play to Investigate how LLMs Present Collective Memory
- 让大模型扮演叙述者,通过谋杀哲学家事件展开互动游戏
- 115份开场文本分析显示不同模型对历史内容呈现差异显著
- 以艺术展形式让公众参与,轻松理解大模型记忆机制
本文报告了艺术研究项目首阶段成果,探讨大型语言模型(LLMs)如何筛选与呈现集体记忆。2025年在维也纳展出两个月的公共装置中,访客可与五种不同模型(ChatGPT GPT 4o 和 GPT 4o mini、Mistral Large、DeepSeek-Chat,以及本地运行的 Llama 3.1)互动,这些模型被设定为叙述者,参与围绕1936年奥地利哲学家莫里茨·施利克谋杀案的角色扮演游戏。研究收集了玩家在游戏中的交互记录及体验后的定性访谈,深入分析玩家反应。定量分析共考察115份由大模型生成的角色扮演开场文本,采用自然语言处理技术进行语义相似度与情感分析。定性反馈揭示出三类用户行为模式,而定量分析表明不同模型在历史内容呈现上存在显著差异。本研究不仅推动对大模型表现力的理解,更提出一种以游戏化方式向公众传播相关认知的新路径。
原文摘要 · Abstract (English)
The paper presents the first results of an artistic research project investigating how Large Language Models (LLMs) curate and present collective memory. In a public installation exhibited during two months in Vienna in 2025, visitors could interact with five different LLMs (ChatGPT with GPT 4o and GPT 4o mini, Mistral Large, DeepSeek-Chat, and a locally run Llama 3.1 model), which were instructed to act as narrators, implementing a role-playing game revolving around the murder of Austrian philosopher Moritz Schlick in 1936. Results of the investigation include protocols of LLM-user interactions during the game and qualitative conversations after the play experience to get insight into the players' reactions to the game. In a quantitative analysis 115 introductory texts for role-playing generated by the LLMs were examined by different methods of natural language processing, including semantic similarity and sentiment analysis. While the qualitative player feedback allowed to distinguish three distinct types of users, the quantitative text analysis showed significant differences between how the different LLMs presented the historical content. Our study thus adds to ongoing efforts to analyse LLM performance, but also suggests a way of how these efforts can be disseminated in a playful way to a general audience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。