arXiv:2409.12380cs.IRcs.AI2024-09被引 2

用子模优化整合网页片段,挖掘更完整的热点主题

Bundle Fragments into a Whole: Mining More Complete Clusters via Submodular Selection of Interesting webpages for Web Topic Detection

  • 先分组后精炼:将碎片化话题聚为粗粒度主题,再用子模方法优化
  • 在两个公开数据集上准确率提升20%,召回率提升10%
  • 适合需要自动发现完整热点的网络趋势分析场景

将有趣网页组织成热点话题是理解多模态网络数据趋势的关键步骤。现有方法先生成大量多粒度话题候选,再通过兴趣度估计识别热点,但这些候选常因特征表示不足和无监督生成而包含大量碎片化信息。本文提出一种捆绑-精炼框架,先将碎片话题合并为粗粒度主题,再基于子模优化进行可扩展的精炼。该方法虽简单却高效,显著优于需精心设计与复杂步骤的传统排序方法。在两个公开数据集上的实验表明,本方法相较最先进模型(即Pang等,2016年的隐式泊松去卷积)分别提升20%准确率和10%召回率。

原文摘要 · Abstract (English)

Organizing interesting webpages into hot topics is one of key steps to understand the trends of multimodal web data. A state-of-the-art solution is firstly to organize webpages into a large volume of multi-granularity topic candidates; hot topics are further identified by estimating their interestingness. However, these topic candidates contain a large number of fragments of hot topics due to both the inefficient feature representations and the unsupervised topic generation. This paper proposes a bundling-refining approach to mine more complete hot topics from fragments. Concretely, the bundling step organizes the fragment topics into coarse topics; next, the refining step proposes a submodular-based method to refine coarse topics in a scalable approach. The propose unconventional method is simple, yet powerful by leveraging submodular optimization, our approach outperforms the traditional ranking methods which involve the careful design and complex steps. Extensive experiments demonstrate that the proposed approach surpasses the state-of-the-art method (i.e., latent Poisson deconvolution Pang et al. (2016)) 20% accuracy and 10% one on two public data sets, respectively.

话题检测子模优化网页聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。