融合多个视觉模型的先验知识,提升大模型空间理解能力
Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding

- 用高效代理生成多源视觉先验,降低推理开销
- 动态融合机制实现上下文感知的先验注入
- 在多个空间推理任务上刷新性能纪录
多模态大语言模型在空间理解方面展现出巨大潜力。现有方法通常从预训练基础模型中提取先验知识以增强模型的空间感知能力。本文首次发现,将不同基础模型引入多模态大模型时,各模型提供的空间先验具有互补性,适用于不同任务。为此,我们提出ViPS框架,旨在充分挖掘来自多种模型的视觉先验以提升空间理解能力。ViPS引入高效先验代理,以极小计算开销生成多重基础先验;并设计动态先验融合机制,实现上下文感知的和谐融合与注入。大量实验表明,ViPS成功协调多样视觉先验,在多个复杂空间推理和3D空间理解基准上达到新最优性能。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have demonstrated substantial promise in spatial understanding. Existing works typically incorporate prior knowledge extracted from a pre-trained foundation model to further enhance the spatial awareness of MLLMs. In this paper, we first reveal that when integrating diverse foundation models into MLLMs, different models provide complementary spatial priors that benefit different tasks. Motivated by this, we propose $\textbf{ViPS}$, a novel multi-model prior framework designed to fully unleash the potential of incorporating multiple $\textbf{Vi}$sual $\textbf{P}$riors from diverse models into MLLMs for $\textbf{S}$patial understanding. Specifically, ViPS introduces an Efficient Prior Proxy to generate multiple foundational priors with minimal inference overhead, and a Dynamic Prior Fusion mechanism to achieve harmonious and context-aware prior fusion and injection from the prior proxies. Extensive experiments demonstrate that ViPS successfully harmonizes diverse visual priors, establishing new state-of-the-art performance across multiple complex spatial reasoning and 3D spatial understanding benchmarks. Project page: https://visual-ai.github.io/vips
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。