用双视觉编码器和新数据集提升模型的空间推理能力
Towards Visuospatial Cognition via Hierarchical Fusion of Visual Experts
- 结合语义与空间结构的双编码器设计,提升空间理解
- 在VSI-Bench上达56.8分,优于更大模型和商用模型
- 适合研究空间推理、多模态认知的学者和开发者
尽管多模态大语言模型在通用视觉-语言任务中表现优异,但对空间布局、关系和动态的视觉空间认知仍面临挑战。现有模型缺乏必要的架构组件和专门训练数据以实现细粒度空间理解。我们提出ViCA2(视觉空间认知助手2),一种新型多模态大语言模型,采用融合SigLIP(语义)与Hiera(空间结构)的双视觉编码器,并引入令牌比例控制机制以提升效率。我们还构建了包含超过32.2万条空间对齐问答对的ViCA-322K大规模数据集,用于定向指令微调。在具有挑战性的VSI-Bench基准测试中,ViCA2-7B模型取得56.8的平均得分,显著超越更大的开源模型(如LLaVA-NeXT-Video-72B,40.9)和领先商用模型(Gemini-1.5 Pro,45.4)。结果表明,该方法能在紧凑模型下实现强视觉空间智能。我们已开源ViCA2、代码库及ViCA-322K数据集,以推动后续研究。
原文摘要 · Abstract (English)
While Multimodal Large Language Models (MLLMs) excel at general vision-language tasks, visuospatial cognition - reasoning about spatial layouts, relations, and dynamics - remains a significant challenge. Existing models often lack the necessary architectural components and specialized training data for fine-grained spatial understanding. We introduce ViCA2 (Visuospatial Cognitive Assistant 2), a novel MLLM designed to enhance spatial reasoning. ViCA2 features a dual vision encoder architecture integrating SigLIP for semantics and Hiera for spatial structure, coupled with a token ratio control mechanism for efficiency. We also developed ViCA-322K, a new large-scale dataset with over 322,000 spatially grounded question-answer pairs for targeted instruction tuning. On the challenging VSI-Bench benchmark, our ViCA2-7B model achieves a state-of-the-art average score of 56.8, significantly surpassing larger open-source models (e.g., LLaVA-NeXT-Video-72B, 40.9) and leading proprietary models (Gemini-1.5 Pro, 45.4). This demonstrates the effectiveness of our approach in achieving strong visuospatial intelligence with a compact model. We release ViCA2, its codebase, and the ViCA-322K dataset to facilitate further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。