用计算方法对比两大幸存者口述史档案,发现结构差异没那么绝对。
The Shape of Testimony: A Scalable Framework for Oral History Archive Comparison

- 通过话题连贯性与提问分布分析,量化口述史的结构程度。
- 1600多份口述史显示两档案存在显著重叠,非简单对立。
- 框架可复用,适合数字口述史与公众参与标注平台。
大屠杀研究者常将美国南加州大学视听基金会的口述史归为结构化、访谈引导型,而耶鲁大学福图诺夫视频档案馆则偏向自由开放型。本研究通过对两个收藏中超过1,600份口述史进行大规模计算分析,利用话语分割、主题建模及大语言模型(LLM)分析,从话题连贯性、访谈互动动态和问题类型分布等维度量化“结构化程度”。结果总体支持既有结构性差异的判断,但也揭示了两馆在单个访谈内及共同叙事模式上的显著重叠,挑战了“结构化对自由式”的简单二分法。本研究不仅重新审视了大屠杀研究中的基础假设,更提供了一个可扩展、可复现的比较语料库分析框架。作为概念验证,该框架为数字口述史、叙事分析及公民科学标注平台的设计提供了广阔应用前景。
原文摘要 · Abstract (English)
Researchers in Holocaust studies have often distinguished between two styles of oral survivor testimony: the USC Shoah Foundation's interviews tend to follow a structured, interviewer-guided format, whereas the Yale Fortunoff Video Archive generally favors a more free-form, open-ended style. This distinction has influenced both scholarly research and the development of later archives. In this study, we critically examine that claim by conducting a large-scale computational analysis of more than 1,600 testimonies from both collections. Leveraging discourse segmentation, topic modeling, and large language model (LLM) based analysis, we quantify the "structuredness" level of testimonies through topic coherence, interviewer-survivor dynamics, and the distribution of question types. Our results generally corroborate the structural differences identified in earlier research, while also revealing significant overlaps between the collections, both within individual interviews and across common narrative patterns. This complicates the simple "structured vs. free-form" dichotomy often applied to these oral histories. Beyond revisiting a foundational claim in Holocaust studies, our work provides a scalable, replicable framework for comparative corpus analysis. As a proof of concept, it suggests broader applications for digital oral history, narrative analysis, and the design of citizen-science annotation platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。