让主题模型懂用户意图,生成更相关、多样的主题。
Human-Centric Topic Modeling with Goal-Prompted Contrastive Learning and Optimal Transport
- 用大模型提取文档中的目标候选,融入对比学习
- 在三个Reddit数据集上主题连贯性与多样性均领先
- 适合需要精准理解用户目标的场景
现有主题建模方法(从LDA到基于神经网络和大模型的方法)主要关注统计连贯性,常产生冗余或偏离用户意图的主题。本文提出人类中心主题建模(Human-TM),将用户提供的目标直接融入建模过程,以生成可解释、多样且目标导向的主题。为此,我们提出GCTM-OT模型:首先利用大模型提示从文档中提取目标候选,再通过最优传输实现语义感知的对比学习来发现主题。在三个公开的Reddit数据集上的实验表明,GCTM-OT在主题连贯性和多样性上优于当前最佳基线,并显著提升与人工提供目标的一致性,为更人性化主题发现系统开辟了道路。
原文摘要 · Abstract (English)
Existing topic modeling methods, from LDA to recent neural and LLM-based approaches, which focus mainly on statistical coherence, often produce redundant or off-target topics that miss the user's underlying intent. We introduce Human-centric Topic Modeling, \emph{Human-TM}), a novel task formulation that integrates a human-provided goal directly into the topic modeling process to produce interpretable, diverse and goal-oriented topics. To tackle this challenge, we propose the \textbf{G}oal-prompted \textbf{C}ontrastive \textbf{T}opic \textbf{M}odel with \textbf{O}ptimal \textbf{T}ransport (GCTM-OT), which first uses LLM-based prompting to extract goal candidates from documents, then incorporates these into semantic-aware contrastive learning via optimal transport for topic discovery. Experimental results on three public subreddit datasets show that GCTM-OT outperforms state-of-the-art baselines in topic coherence and diversity while significantly improving alignment with human-provided goals, paving the way for more human-centric topic discovery systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。