arXiv:2608.23284cs.CL2026-08中稿 · CIKM 2026

提出共享主题空间+残差适配的新框架,实现跨语料库长期主题演化稳定对比。

Dynamic Topic Modeling for Cross-Corpus Temporal Analysis

论文配图:Dynamic Topic Modeling for Cross-Corpus Temporal Analysis
图 1 · 摘自论文原文
  • 先构建共享动态主题空间,再用残差项适配各语料特异性。
  • 跨语料主题轨迹匹配率97.5%,远超全微调的17.9%。
  • 适合需要长期跨语料主题追踪的研究者使用。

动态嵌入主题模型(D-ETM)能解释语义随时间演变,但跨语料比较困难,因主题常独立学习且仅在训练后对齐,无法保证跨语料和时间的主题对应稳定性。为此,我们提出一种D-ETM框架:首先在合并的多语料集合上学习一个共同的动态主题空间(称作共享骨干),然后在冻结的骨干基础上引入语料特定的残差适配,不创建独立的潜在主题空间。该设计保留了统一的主题索引,便于跨语料比较,同时允许各语料保留词汇特异性。我们在三个历时结构语料库(涵盖97年)上评估:美国历史英语语料库、哈佛商业评论、国际劳工评论。残差适配在提升语料特异性拟合的同时,保持相同索引的跨语料主题轨迹,其轨迹检索准确率@1达97.5±0.7%,显著优于从同一骨干全微调的17.9±1.1%,也优于独立训练后采用匈牙利匹配的方案。结果表明,将主题对齐纳入建模可支持更稳定的跨时间、跨语料比较,同时保留语料特异性词汇变化。

原文摘要 · Abstract (English)

Dynamic Embedded Topic Models (D-ETM) provide an interpretable framework for modeling temporal semantic evolution, but cross-corpus comparison remains difficult because topics are often learned independently and aligned only after training, a process that does not guarantee stable topic correspondence across corpora and time. To address this problem, we propose a D-ETM framework that first learns a common dynamic topic space over a merged multi-corpus collection, which we call the shared backbone, then introduces corpus-specific residual adaptation around the frozen backbone without creating separate latent topic spaces. This design preserves a shared topic index for cross-corpus comparison while allowing each corpus to specialize lexically. We evaluate the framework on three temporally structured corpora spanning 97 years: the Corpus of Historical American English, Harvard Business Review, and International Labour Review. Residual adaptation improves corpus-specific fit relative to the shared backbone while preserving the same-index cross-corpus topic trajectories, achieving substantially stronger alignment than full fine-tuning from the same backbone, with $97.5 \pm 0.7\%$ versus $17.9 \pm 1.1\%$ trajectory Retrieval@1, as well as stronger alignment than independent training with post-hoc Hungarian matching. These results suggest that incorporating topic alignment into the model can support more stable over-time cross-corpus comparisons while retaining corpus-specific lexical variation.

主题建模动态主题跨语料分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。