用大模型分析斯洛文尼亚百年报纸,揭示集体身份与政治立场的演变。
Approaches to Analysing Historical Newspapers Using LLMs
- 结合主题建模与大模型情感分析,识别报纸核心议题与意识形态差异。
- 在手写体历史文本中,选定GaMS3-12B-Instruct模型实现高精度情感分类。
- 通过实体图谱揭示群体、地域与政治认同之间的复杂关系,适合数字人文研究者。
本研究基于斯洛文尼亚历史报纸语料库sPeriodika中的《Slovenec》和《Slovenski narod》,采用主题建模、大语言模型(LLM)层面的情感分析、实体图谱可视化与定性话语分析相结合的方法,考察20世纪初公共话语中集体身份、政治倾向与民族归属的表达方式。利用BERTopic识别出主要主题模式,显示两份报纸在保守天主教与自由进步立场上存在明显分歧。评估四种指令遵循型大模型在光学字符识别降质的历史斯洛文尼亚语上的表现后,选定经本地化适配的GaMS3-12B-Instruct模型作为大规模应用的最佳选择,但发现其对中性情绪识别优于正向或负向情绪。应用于全数据集时,模型揭示了不同群体在叙述中呈现的显著差异:部分群体多出现在中性描述中,另一些则更常与评价性或冲突性话语关联。随后构建命名实体图谱,探究集体身份与地理空间的关系,并结合定量网络分析与批判话语分析,聚焦历史政治与社会经济身份的交织演进。研究表明,将可扩展计算方法与批判性解读结合,能有效支持噪声较大的历史报纸数据的数字人文研究。
原文摘要 · Abstract (English)
This study presents a computational analysis of the Slovene historical newspapers \textit{Slovenec} and \textit{Slovenski narod} from the sPeriodika corpus, combining topic modelling, large language model (LLM)-based aspect-level sentiment analysis, entity-graph visualisation, and qualitative discourse analysis to examine how collective identities, political orientations, and national belonging were represented in public discourse at the turn of the twentieth century. Using BERTopic, we identify major thematic patterns and show both shared concerns and clear ideological differences between the two newspapers, reflecting their conservative-Catholic and liberal-progressive orientations. We further evaluate four instruction-following LLMs for targeted sentiment classification in OCR-degraded historical Slovene and select the Slovene-adapted GaMS3-12B-Instruct model as the most suitable for large-scale application, while also documenting important limitations, particularly its stronger performance on neutral sentiment than on positive or negative sentiment. Applied at dataset scale, the model reveals meaningful variation in the portrayal of collective identities, with some groups appearing predominantly in neutral descriptive contexts and others more often in evaluative or conflict-related discourse. We then create NER graphs to explore the relationships between collective identities and places. We apply a mixed methods approach to analyse the named entity graphs, combining quantitative network analysis with critical discourse analysis. The investigation focuses on the emergence and development of intertwined historical political and socionomic identities. Overall, the study demonstrates the value of combining scalable computational methods with critical interpretation to support digital humanities research on noisy historical newspaper data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。