用AI从报纸档案中自动提取历史话题,看清核能讨论的演变
Automating Historical Insight Extraction from Large-Scale Newspaper Archives via Neural Topic Modeling
- 用BERTopic模型分析1955-2018年报纸文本,捕捉动态话题变化
- 发现核能与核武器话题在公众讨论中的共现关系及重要性变迁
- 适合历史学、社会科学研究者,尤其关注舆论演变与技术议题
从大规模非结构化历史报纸档案中提取连贯且可理解的主题面临巨大挑战,主要源于话题演变、光学字符识别(OCR)噪声以及文本体量庞大。传统主题模型如隐含狄利克雷分配(LDA)难以捕捉历史文本中话语的复杂性和动态性。为此,本文采用BERTopic这一神经主题建模方法,利用基于Transformer的嵌入表示进行主题提取与分类。尽管该方法日益流行,但在历史研究中仍应用不足。研究聚焦1955至2018年间关于核能与核安全的报道,分析不同主题在语料库中的分布及其随时间的演变,揭示长期趋势与公共话语的转变。结果表明,该方法在可扩展性与上下文敏感性方面优于传统方法,能更准确地揭示核能与核武器相关话题的共现模式及其重要性变迁。研究为历史学、核科学与社会科学提供新洞见,同时指出当前局限并提出未来方向。
原文摘要 · Abstract (English)
Extracting coherent and human-understandable themes from large collections of unstructured historical newspaper archives presents significant challenges due to topic evolution, Optical Character Recognition (OCR) noise, and the sheer volume of text. Traditional topic-modeling methods, such as Latent Dirichlet Allocation (LDA), often fall short in capturing the complexity and dynamic nature of discourse in historical texts. To address these limitations, we employ BERTopic. This neural topic-modeling approach leverages transformerbased embeddings to extract and classify topics, which, despite its growing popularity, still remains underused in historical research. Our study focuses on articles published between 1955 and 2018, specifically examining discourse on nuclear power and nuclear safety. We analyze various topic distributions across the corpus and trace their temporal evolution to uncover long-term trends and shifts in public discourse. This enables us to more accurately explore patterns in public discourse, including the co-occurrence of themes related to nuclear power and nuclear weapons and their shifts in topic importance over time. Our study demonstrates the scalability and contextual sensitivity of BERTopic as an alternative to traditional approaches, offering richer insights into historical discourses extracted from newspaper archives. These findings contribute to historical, nuclear, and social-science research while reflecting on current limitations and proposing potential directions for future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。