arXiv:2506.11394cs.CVcs.AI2025-06

用动态空间塔增强模型对图像空间关系的理解能力。

Dynamic Double Space Tower

  • 设计四层动态双向空间塔,模拟人类整体视觉感知。
  • 在空间关系问答任务上达顶尖性能,仅用30亿参数。
  • 可无缝接入任意多模态模型,提升推理效率。

视觉问答(VQA)任务需要同时理解图像内容与问题语义。现有方法常因跨模态交互不足及难以捕捉图像中实体的空间关系而受限。本文提出一种新型动态双向空间塔结构,共分四层,依据人类格式塔视觉原理观察图像,自然为实体间空间组织提供强大结构先验,使模型从盲目搜索像素关系转向基于有意义感知单元的判断,实现从“看图”到“感知并组织图像内容”的转变。大量实验表明,该模块可嵌入任意多模态模型并取得先进效果,尤其在空间关系问答数据集上表现突出。基于该方法训练的多模态视觉问答模型July,在仅30亿参数下达到当前最优水平。

原文摘要 · Abstract (English)

The Visual Question Answering (VQA) task requires the simultaneous understanding of image content and question semantics. However, existing methods often have difficulty handling complex reasoning scenarios due to insufficient cross-modal interaction and capturing the entity spatial relationships in the image.\cite{huang2023adaptive}\cite{liu2021comparing}\cite{guibas2021adaptive}\cite{zhang2022vsa}We studied a brand-new approach to replace the attention mechanism in order to enhance the reasoning ability of the model and its understanding of spatial relationships.Specifically, we propose a dynamic bidirectional spatial tower, which is divided into four layers to observe the image according to the principle of human gestalt vision. This naturally provides a powerful structural prior for the spatial organization between entities, enabling the model to no longer blindly search for relationships between pixels but make judgments based on more meaningful perceptual units. Change from "seeing images" to "perceiving and organizing image content".A large number of experiments have shown that our module can be used in any other multimodal model and achieve advanced results, demonstrating its potential in spatial relationship processing.Meanwhile, the multimodal visual question-answering model July trained by our method has achieved state-of-the-art results with only 3B parameters, especially on the question-answering dataset of spatial relations.

视觉问答空间关系多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。