arXiv:2501.00330cs.CLcs.AI2025-01被引 5

用多模态大模型挖掘实体隐含语义,提升实体扩展效果。

Exploring the Implicit Semantic Ability of Multimodal Large Language Models: A Pilot Study on Entity Set Expansion

  • 设计列表排序方法LUSAR,将局部得分映射为全局排序。
  • 在多模态实体扩展任务中显著提升大模型表现。
  • 首次将生成式多模态大模型用于实体集扩展,适合跨模态研究者。

多模态大语言模型(MLLMs)的快速发展推动了实际应用中多项任务的进展。然而,现有大模型在提取隐含语义信息方面仍存在局限。本文将MLLMs应用于多模态实体集扩展(MESE)任务,旨在基于少量种子实体,扩展出同语义类的新实体,并为每个实体提供多模态信息。通过MESE任务,我们探索了MLLM在实体层面理解隐含语义的能力,提出一种列表排序方法LUSAR,将局部得分映射为全局排名。实验表明,LUSAR显著提升了MLLM在MESE任务上的性能,这是首个将生成式多模态大模型应用于实体集扩展的工作,拓展了列表排序方法的应用边界。

原文摘要 · Abstract (English)

The rapid development of multimodal large language models (MLLMs) has brought significant improvements to a wide range of tasks in real-world applications. However, LLMs still exhibit certain limitations in extracting implicit semantic information. In this paper, we apply MLLMs to the Multi-modal Entity Set Expansion (MESE) task, which aims to expand a handful of seed entities with new entities belonging to the same semantic class, and multi-modal information is provided with each entity. We explore the capabilities of MLLMs to understand implicit semantic information at the entity-level granularity through the MESE task, introducing a listwise ranking method LUSAR that maps local scores to global rankings. Our LUSAR demonstrates significant improvements in MLLM's performance on the MESE task, marking the first use of generative MLLM for ESE tasks and extending the applicability of listwise ranking.

多模态实体扩展大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。