arXiv:2606.14348physics.soc-phcond-mat.stat-mech2026-06

用复杂系统方法分析60万篇意大利报纸,发现重大历史转折点

Detecting Historical Turning Points in Italian Media: A Complex Systems Approach to a Diachronic News Corpus

论文配图:Detecting Historical Turning Points in Italian Media: A Complex Systems Approach to a Diachronic News Corpus
图 1 · 摘自论文原文
  • 构建1985-2000年60万篇《共和报》语料库,结合词法与语义分析
  • 自动识别出意大利第一共和国向第二共和国过渡等关键转折期
  • 无需人工标注,适合对媒体演化、历史变迁感兴趣的学者

大规模文本语料库的出现为基于自然语言处理(NLP)的数据驱动历史分析提供了新可能。然而,具有历史意义且覆盖前数字时代的历时语料库仍稀少且不完整。本文基于意大利《共和报》1985年1月1日至2000年12月31日间约60万篇报道,构建历时语料库,涵盖意大利及全球重大政治、社会与地缘政治变革时期。通过NLP技术在词汇与语义层面分析文本,并应用复杂系统与统计物理工具,追踪媒体话语随时间的演变。结果可自动检测出从第一共和国到第二共和国的转型、海湾战争、科索沃战争等关键历史转折点,无需依赖人工标注。研究证明,计算语言学与复杂系统思想结合,能为历史变迁提供新的量化洞察,开辟了通过大规模文本数据研究媒体与社会动态的新路径。

原文摘要 · Abstract (English)

The increasing availability of large-scale textual corpora has opened new possibilities for data-driven, quantitative approaches to historical analysis using Natural Language Processing (NLP). However, diachronic corpora with historical relevance from the pre-digital era remain scarce and often incomplete. We present a quantitative approach to historical analysis based on the reconstruction and exploration of a diachronic corpus of around 600,000 articles from the Italian newspaper "La Repubblica", covering all the articles published from the 1st of January 1985 to the 31st of December 2000 - a period of major political, social, and geopolitical change in Italy and globally. Using NLP techniques, we analyze the text at both lexical and semantic levels; we then apply tools from complex systems and statistical physics to trace shifts in media discourse over time. This allows us to detect key transition periods, such as the transition from the First Republic to the Second Republic in Italy, or major international conflicts like the Gulf War or the Kosovo War, without relying on prior labeling. The results show how combining computational linguistics with ideas from complex systems can offer new quantitative insight into historical changes, opening up new paths for studying the dynamics of media and society through large-scale textual data.

历史分析复杂系统文本挖掘媒体演化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。