无需训练即可补全缺失模态,且在跨域场景下表现更优。
Knowledge Bridger: Towards Training-free Missing Modality Completion
- 利用大模型构建知识图谱,实现跨模态信息连接与推理。
- 在通用和医疗领域均超越现有方法,尤其在跨域场景下优势明显。
- 无需训练、适配性强,适合实际应用中的快速部署。
以往的缺失模态补全方法依赖精心设计的融合技术与完整数据上的大量预训练,限制了其在域外(OOD)场景下的泛化能力。本文提出新挑战:能否构建一个资源高效且对域外泛化鲁棒的缺失模态补全模型?为此,我们提出一种无需训练的框架,基于大模型(LMM)实现缺失模态补全。所提方法“Knowledge Bridger”具有模态无关性,整合生成与排序机制。通过定义特定领域先验,自动从已有模态中提取结构化信息构建知识图谱,再通过大模型将缺失模态生成与排序模块连接,实现高质量补全。在通用与医疗领域的实验表明,该方法持续优于对比方法,尤其在域外泛化上表现突出。此外,基于知识的生成与排序优于直接使用大模型的方法,为其他领域应用提供有益启示。
原文摘要 · Abstract (English)
Previous successful approaches to missing modality completion rely on carefully designed fusion techniques and extensive pre-training on complete data, which can limit their generalizability in out-of-domain (OOD) scenarios. In this study, we pose a new challenge: can we develop a missing modality completion model that is both resource-efficient and robust to OOD generalization? To address this, we present a training-free framework for missing modality completion that leverages large multimodal models (LMMs). Our approach, termed the "Knowledge Bridger", is modality-agnostic and integrates generation and ranking of missing modalities. By defining domain-specific priors, our method automatically extracts structured information from available modalities to construct knowledge graphs. These extracted graphs connect the missing modality generation and ranking modules through the LMM, resulting in high-quality imputations of missing modalities. Experimental results across both general and medical domains show that our approach consistently outperforms competing methods, including in OOD generalization. Additionally, our knowledge-driven generation and ranking techniques demonstrate superiority over variants that directly employ LMMs for generation and ranking, offering insights that may be valuable for applications in other domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。