arXiv:2506.00277cs.CLcs.AI2025-06ACL被引 4

用分层嵌入实现多语言新闻聚类,高效又可解释。

Hierarchical Level-Wise News Article Clustering via Multilingual Matryoshka Embeddings

  • 通过嵌入维度分级判断新闻相似性,支持细粒度到粗粒度聚类。
  • 在SemEval 2022任务中达0.816的皮尔逊相关系数,性能领先。
  • 适合处理多语言新闻数据,尤其适合需可解释聚类的场景。

上下文大模型嵌入被广泛用于主题建模与聚类,但现有方法常面临扩展性差、相似度度量不透明、多语言表现不佳等问题。本文提出一种新型、可扩展、可解释、分层且支持多语言的新闻文章与社交媒体数据聚类方法。首先,训练多语言马特罗什卡嵌入,通过分析嵌入向量的不同维度子集,实现不同粒度的故事相似性判断。该模型在SemEval 2022任务8测试集上取得0.816的皮尔逊相关系数(Pearson $ρ$),达到当前最佳水平。训练完成后,设计一种高效分层聚类算法,利用马特罗什卡嵌入的层级特性,识别出独立新闻事件、叙事线索与核心主题。最后,通过真实新闻数据集验证了该方法在发现和聚类故事、叙事及宏观主题方面的有效性。

原文摘要 · Abstract (English)

Contextual large language model embeddings are increasingly utilized for topic modeling and clustering. However, current methods often scale poorly, rely on opaque similarity metrics, and struggle in multilingual settings. In this work, we present a novel, scalable, interpretable, hierarchical, and multilingual approach to clustering news articles and social media data. To do this, we first train multilingual Matryoshka embeddings that can determine story similarity at varying levels of granularity based on which subset of the dimensions of the embeddings is examined. This embedding model achieves state-of-the-art performance on the SemEval 2022 Task 8 test dataset (Pearson $ρ$ = 0.816). Once trained, we develop an efficient hierarchical clustering algorithm that leverages the hierarchical nature of Matryoshka embeddings to identify unique news stories, narratives, and themes. We conclude by illustrating how our approach can identify and cluster stories, narratives, and overarching themes within real-world news datasets.

新闻聚类多语言分层嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。