arXiv:2605.03299cs.CL2026-05ACL被引 1

用大模型提升跨语言主题模型,更准更省资源。

LLM-XTM: Enhancing Cross-Lingual Topic Models with Large Language Models

论文配图:LLM-XTM: Enhancing Cross-Lingual Topic Models with Large Language Models
图 1 · 摘自论文原文
  • 用大模型引导主题优化,黑盒操作无须词频概率
  • 主题一致性与跨语对齐效果更好,减少对双语词典依赖
  • 适合需要低成本高精度跨语言分析的研究者

跨语言主题建模旨在发现多语言间的共享语义结构,但现有方法依赖稀疏的双语资源,常产生不连贯或弱对齐的主题。近期基于大模型的改进虽提升了可解释性,但代价高昂、仅限文档级处理,且易产生幻觉;以往白盒方法还需难以获取的词元概率。我们提出 LLM-XTM 框架,融合大模型引导的主题精炼与自一致不确定性量化,实现黑盒、稳定且可扩展的跨语言主题模型增强。在多语言语料上的实验表明,LLM-XTM 在提升主题一致性与跨语对齐效果的同时,降低了对双语词典和昂贵大模型调用的依赖。

原文摘要 · Abstract (English)

Cross-lingual topic modeling aims to discover shared semantic structures across languages, yet existing models depend on sparse bilingual resources and often yield incoherent or weakly aligned topics. Recent LLM-based refinements improve interpretability but are costly, document-level, and prone to hallucination, with prior white-box approaches requiring inaccessible token probabilities. We propose LLM-XTM, a framework that integrates LLM-guided topic refinement with self-consistency uncertainty quantification, enabling black-box, stable, and scalable enhancement of cross-lingual topic models. Experiments on multilingual corpora show that LLM-XTM achieves superior topic coherence and alignment while reducing reliance on bilingual dictionaries and expensive LLM calls.

主题模型跨语言大模型生成式AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。