通过改进维度消除语言偏差,提升跨语言主题模型准确性
Refining Dimensions for Improving Clustering-based Cross-lingual Topic Models
- 用SVD方法精炼上下文向量维度,消除多语言模型带来的语言特异性干扰
- 在三个数据集上表现优于现有顶尖跨语言主题模型
- 适合需要跨语言文本分析的研究者,尤其关注主题建模的场景
基于聚类的主题模型在单语主题识别中表现良好,通过引入聚类上下文表示的流程实现。然而,由于多语言语言模型生成的语言依赖维度(LDDs),该流程在跨语言主题识别中表现不佳。为此,本文提出一种基于SVD的维度精炼组件,嵌入聚类主题模型流程中,有效抵消LDDs的负面影响,使模型能更准确地识别跨语言主题。在三个数据集上的实验表明,加入该组件后的流程普遍优于其他最先进的跨语言主题模型。
原文摘要 · Abstract (English)
Recent works in clustering-based topic models perform well in monolingual topic identification by introducing a pipeline to cluster the contextualized representations. However, the pipeline is suboptimal in identifying topics across languages due to the presence of language-dependent dimensions (LDDs) generated by multilingual language models. To address this issue, we introduce a novel, SVD-based dimension refinement component into the pipeline of the clustering-based topic model. This component effectively neutralizes the negative impact of LDDs, enabling the model to accurately identify topics across languages. Our experiments on three datasets demonstrate that the updated pipeline with the dimension refinement component generally outperforms other state-of-the-art cross-lingual topic models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。