arXiv:2510.13542cs.LG2025-10

用原型网络提升医疗文献少样本主题建模效果

ProtoTopic: Prototypical Network for Few-Shot Medical Topic Modeling

  • 基于原型网络计算文档与主题原型距离,实现少样本学习
  • 在医疗文献上生成主题的连贯性与多样性优于两个基线模型
  • 适合数据稀缺的医学文本分析场景

主题建模是分析大量文本(尤其是学术论文)的有力工具。尽管已有多种主题建模方法,但在医疗文本上表现不佳,原因可能是某些主题可用文档数量过少。本文提出ProtoTopic,一种基于原型网络的主题模型,用于医学论文摘要的主题生成。原型网络通过计算输入样本与一组原型表示的距离进行预测,具有高效且可解释的特点,特别适用于低数据或少样本学习场景。实验表明,与文献中的两种基线模型相比,ProtoTopic在主题连贯性和多样性上均有提升,证明其能在数据有限的情况下生成具有医学相关性的主题。

原文摘要 · Abstract (English)

Topic modeling is a useful tool for analyzing large corpora of written documents, particularly academic papers. Despite a wide variety of proposed topic modeling techniques, these techniques do not perform well when applied to medical texts. This can be due to the low number of documents available for some topics in the healthcare domain. In this paper, we propose ProtoTopic, a prototypical network-based topic model used for topic generation for a set of medical paper abstracts. Prototypical networks are efficient, explainable models that make predictions by computing distances between input datapoints and a set of prototype representations, making them particularly effective in low-data or few-shot learning scenarios. With ProtoTopic, we demonstrate improved topic coherence and diversity compared to two topic modeling baselines used in the literature, demonstrating the ability of our model to generate medically relevant topics even with limited data.

主题建模少样本学习医疗文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。