用视觉语言模型增强多模态主题建模,提升文档理解与跨模态一致性。
CEMTM: Contextual Embedding-based Multimodal Topic Modeling
- 基于微调的视觉语言模型生成上下文嵌入,融合图文信息。
- 在六个数据集上平均大模型评分达2.61,优于现有方法。
- 支持多图输入不重复编码,适合科学论文等复杂场景。
我们提出CEMTM,一种基于上下文嵌入的多模态主题模型,用于从包含文本和图像的短长文档中推断连贯且可解释的主题结构。CEMTM利用微调的大规模视觉语言模型(LVLMs)获取上下文嵌入,并采用分布注意力机制加权词级贡献以进行主题推断。通过重建目标,使基于主题的表示与文档嵌入对齐,促进跨模态语义一致性。与现有方法不同,CEMTM可在单篇文档中处理多个图像而无需重复编码,并通过显式的词-主题与文档-主题分布保持可解释性。在六个多模态基准上的广泛实验表明,CEMTM始终优于单模态和多模态基线,平均大模型评分达到2.61。进一步分析显示其在下游少样本检索任务中的有效性,以及在科学文章等复杂领域捕捉视觉语义的能力。
原文摘要 · Abstract (English)
We introduce CEMTM, a context-enhanced multimodal topic model designed to infer coherent and interpretable topic structures from both short and long documents containing text and images. CEMTM builds on fine-tuned large vision language models (LVLMs) to obtain contextualized embeddings, and employs a distributional attention mechanism to weight token-level contributions to topic inference. A reconstruction objective aligns topic-based representations with the document embedding, encouraging semantic consistency across modalities. Unlike existing approaches, CEMTM can process multiple images per document without repeated encoding and maintains interpretability through explicit word-topic and document-topic distributions. Extensive experiments on six multimodal benchmarks show that CEMTM consistently outperforms unimodal and multimodal baselines, achieving a remarkable average LLM score of 2.61. Further analysis shows its effectiveness in downstream few-shot retrieval and its ability to capture visually grounded semantics in complex domains such as scientific articles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。