检测大模型是否在无意中传播大屠杀否认主义,揭示其记忆建构风险。
From prosthetic memory to prosthetic denial: Auditing whether large language models are prone to mass atrocity denialism
- 对比五款大模型在四起历史事件中的回应,测试其对否认言论的敏感度。
- 对广为人知的犹太人大屠杀回应准确,但对柬埔寨大屠杀等事件易受否认叙事影响。
- 提醒公众警惕大模型可能成为历史记忆扭曲的传播工具,尤其在数据稀疏领域。
大语言模型(LLMs)的普及可能影响历史叙事的传播与认知。本研究探讨了生成式AI在大规模暴行记忆表征上的影响,检验其是否促成‘代理记忆’(mediated experiences of historical events),或我们所称的‘代理否认’——即通过人工智能中介对暴行记忆的抹除或歪曲。我们指出,LLMs作为接口可激发代理记忆,成为记忆传递的体验场所,但也存在否认主义风险,尤其是在其输出与争议性或修正主义叙事一致时。为实证评估这些风险,我们对五款模型(Claude、GPT、Llama、Mixtral、Gemini)在四个历史案例(霍洛多莫尔、大屠杀、柬埔寨大屠杀、卢旺达图西族大屠杀)中进行比较审计,每个案例均以英文及对应相关语言(乌克兰语、德语、高棉语、法语)提出常见否认主张问题。结果表明,尽管对广泛记录的大屠杀如大屠杀回应普遍准确,但在更少被关注的案例如柬埔寨大屠杀中,出现显著不一致且易受否认框架影响。差异凸显训练数据可用性与模型概率响应对记忆完整性的深刻影响。结论认为,尽管LLMs扩展了代理记忆概念,但未经监管的使用可能强化历史否认主义,引发数字记忆保存的伦理关切,并挑战技术原本促进记忆的积极价值。
原文摘要 · Abstract (English)
The proliferation of large language models (LLMs) can influence how historical narratives are disseminated and perceived. This study explores the implications of LLMs' responses on the representation of mass atrocity memory, examining whether generative AI systems contribute to prosthetic memory, i.e., mediated experiences of historical events, or to what we term "prosthetic denial," the AI-mediated erasure or distortion of atrocity memories. We argue that LLMs function as interfaces that can elicit prosthetic memories and, therefore, act as experiential sites for memory transmission, but also introduce risks of denialism, particularly when their outputs align with contested or revisionist narratives. To empirically assess these risks, we conducted a comparative audit of five LLMs (Claude, GPT, Llama, Mixtral, and Gemini) across four historical case studies: the Holodomor, the Holocaust, the Cambodian Genocide, and the genocide against the Tutsis in Rwanda. Each model was prompted with questions addressing common denialist claims in English and an alternative language relevant to each case (Ukrainian, German, Khmer, and French). Our findings reveal that while LLMs generally produce accurate responses for widely documented events like the Holocaust, significant inconsistencies and susceptibility to denialist framings are observed for more underrepresented cases like the Cambodian Genocide. The disparities highlight the influence of training data availability and the probabilistic nature of LLM responses on memory integrity. We conclude that while LLMs extend the concept of prosthetic memory, their unmoderated use risks reinforcing historical denialism, raising ethical concerns for (digital) memory preservation, and potentially challenging the advantageous role of technology associated with the original values of prosthetic memory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。