构建可复现的跨源主题空间,实现不同媒体主题的直接对比
A Shared IPTC Topic Space for Cross-Source Topic Modelling
- 基于IPTC主题分类体系,用引导式BERTopic生成统一主题空间
- 在纽约时报2011数据集上,主题覆盖率达94个一级主题中的高比例
- 支持严格阈值下的稳定降维,适合跨媒体主题分析与对比研究
跨媒体主题注意力比较受制于核心建模问题:各自语料独立训练的主题模型产生不可直接对齐的语料特异性主题空间。本文提出一个可复现框架,将多个语料置于由分类体系定义的共享主题空间中。通过引导式BERTopic发现主题,利用加权关键词与目标中心点,将主题映射到94个IPTC Media Topics一级主题,并依据最大相似性规则合并为17个父级主题。该框架在控制性的《纽约时报》2011语料上经过多轮筛选:广义模型筛选、聚焦映射优化、严格决赛对比、目标构造消融实验及阈值校准。结果显示,在更严格的分配阈值下,引导式方法仍保持显著更高的映射覆盖率;父级增强的目标构造提升了覆盖率与父类一致性;随着阈值收紧,覆盖率呈渐进下降而非骤降。贡献在于提供一种外部锚定的共享主题空间构建方法,实现可复现的跨源主题比较。
原文摘要 · Abstract (English)
Comparing topic attention across different media is hindered by a fundamental modelling problem: topic models fitted separately to each corpus produce corpus-specific topic spaces that cannot be aligned directly. This paper presents a reproducible framework that places corpora in a single shared topic space defined by a taxonomy. Discovered topics are obtained with guided BERTopic, scored against the ninety-four IPTC Media Topics' taxonomy topics (level-1) through weighted keyword and target centroids, and then collapsed upward to seventeen IPTC parent topics by a maximum-similarity rule. The framework was developed and selected on a controlled New York Times 2011 corpus through a narrowing sequence: a broad model screen, a focused mapping refinement, a strict finalist comparison, a target-construction ablation, and a threshold calibration. In this corpus, the guided family retained substantially stronger mapped coverage than a zero-shot benchmark under stricter assignment thresholds, a parent-enriched target construction improved both coverage and parent consistency, and coverage declined gradually rather than collapsing as the assignment threshold was tightened. The contribution is an externally anchored method for constructing a shared topic space that enables reproducible cross-source topic comparison.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。