用多语言模型分析斯里兰卡议会辩论,发现30个跨语主题并追踪重大事件。
Trilingual Topic Modeling of Sri Lankan Parliamentary Debates

- 基于大模型提取复杂排版文本,再用多语言嵌入聚类建模主题。
- 从2017-2026年1.95万篇演讲中识别出30个宏观主题,纯度达0.673。
- 可处理混合语言和黏着形态,适合研究多语社会的政治话语。
斯里兰卡议会辩论(Hansards)构成包含僧伽罗语、泰米尔语和英语的三语语料库,包含代码混杂内容,但因布局复杂的PDF格式、多语言字符集及黏着形态,难以被标准NLP工具处理。本文提出端到端框架,先通过大模型进行文本提取,再经多语言嵌入与密度聚类实现主题建模。进一步引入混合语义-词汇扩展方法BiTopic以提升可解释性,并恢复原本被视作噪声的语篇。应用于2017至2026年间共19,553篇演讲,该流程成功识别出30个宏观主题,聚类纯度(BCP)达0.673,其时间演变轨迹在无监督下与2019年复活节爆炸案及2022年经济危机等重大事件高度吻合。传统LDA因跨语言碎片化而失效,而本方法无需标注即可在三语间捕捉主题结构。
原文摘要 · Abstract (English)
Sri Lankan parliamentary debates (Hansards) constitute a trilingual corpus of speeches in Sinhala, Tamil, and English, including code-mixed content, yet remain inaccessible to standard NLP pipelines due to layout-complex PDFs, multilingual scripts, and agglutinative morphology. We present an end-to-end framework that addresses these challenges through LLM-based text extraction followed by a multilingual embedding and density-based clustering pipeline for topic modeling. A hybrid semantic-lexical extension, BiTopic, is further explored to improve interpretability and recover speeches otherwise discarded as noise. Applied to 19,553 speeches spanning 2017-2026, the pipeline recovers 30 macro-topics achieving a cluster purity (BCP) of 0.673, whose temporal trajectories align unsupervised with major national events including the 2019 Easter Sunday attacks and the 2022 economic crisis. Traditional LDA fails on this corpus due to cross-lingual fragmentation, whereas the proposed approach successfully identifies thematic structure across all three languages without supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。