用视觉语言模型增强多模态知识图谱的对齐与推理能力
VL-KGE: Vision-Language Models Meet Knowledge Graph Embeddings
- 融合视觉语言模型实现跨模态对齐,统一学习多模态实体表示
- 在WN9-IMG和两个艺术知识图谱上,链接预测性能超越传统方法
- 适合需要处理异构多模态数据的智能系统研发者
现实世界的多模态知识图谱(MKG)具有内在异质性,实体关联多种模态。传统知识图谱嵌入(KGE)擅长学习实体与关系的连续表示,但通常针对单模态场景设计。近期方法虽扩展至多模态,但仍受限于模态孤立处理,导致跨模态对齐弱,且依赖实体模态均等可用的简化假设。视觉语言模型(VLMs)能有效将多样模态映射到共享嵌入空间。我们提出视觉语言知识图谱嵌入(VL-KGE),将VLM的跨模态对齐能力与结构化关系建模结合,学习知识图谱的统一多模态表示。在WN9-IMG及两个新构建的细粒度艺术知识图谱WikiArt-MKG-v1和WikiArt-MKG-v2上的实验表明,VL-KGE在链接预测任务中持续优于传统单模态与多模态KGE方法。结果凸显了VLM在多模态KGE中的价值,支持在大规模异构知识图谱上进行更鲁棒、结构化的推理。
原文摘要 · Abstract (English)
Real-world multimodal knowledge graphs (MKGs) are inherently heterogeneous, modeling entities that are associated with diverse modalities. Traditional knowledge graph embedding (KGE) methods excel at learning continuous representations of entities and relations, yet they are typically designed for unimodal settings. Recent approaches extend KGE to multimodal settings but remain constrained, often processing modalities in isolation, resulting in weak cross-modal alignment, and relying on simplistic assumptions such as uniform modality availability across entities. Vision-Language Models (VLMs) offer a powerful way to align diverse modalities within a shared embedding space. We propose Vision-Language Knowledge Graph Embeddings (VL-KGE), a framework that integrates cross-modal alignment from VLMs with structured relational modeling to learn unified multimodal representations of knowledge graphs. Experiments on WN9-IMG and two novel fine art MKGs, WikiArt-MKG-v1 and WikiArt-MKG-v2, demonstrate that VL-KGE consistently improves over traditional unimodal and multimodal KGE methods in link prediction tasks. Our results highlight the value of VLMs for multimodal KGE, enabling more robust and structured reasoning over large-scale heterogeneous knowledge graphs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。