通过多任务学习增强视觉空间推理能力,提升模型对3D场景的理解。
Spatial-ViLT: Enhancing Visual Spatial Reasoning through Multi-Task Learning
- 引入深度图、3D坐标等空间特征,结合多任务学习优化视觉语言表示。
- 在VSR数据集上实现当前最优性能,尤其在方向、拓扑和距离关系推理上表现突出。
- 适合需要精准空间理解的自动驾驶、机器人导航等实际应用场景。
视觉语言模型(VLMs)在多模态推理方面已取得进展,但在三维场景和复杂物体配置下的空间推理仍存在挑战。为此,我们提出SpatialViLT,一种通过多任务学习框架整合深度图、3D坐标和边缘图等空间特征的增强型视觉语言模型。该方法丰富了多模态嵌入的空间感知能力。我们设计了两种变体:SpatialViLT(处理完整物体区域)与MaskedSpatialViLT(聚焦掩码物体区域),并进一步提出SpatialEnsemble融合二者,达到当前最佳准确率。实验表明,该模型在视觉空间推理(VSR)数据集上的方向性、拓扑性和邻近关系推理任务中均表现优异。本工作推动了AI系统空间智能的发展,对实现高级多模态理解及现实应用具有重要意义。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have advanced multimodal reasoning but still face challenges in spatial reasoning for 3D scenes and complex object configurations. To address this, we introduce SpatialViLT, an enhanced VLM that integrates spatial features like depth maps, 3D coordinates, and edge maps through a multi-task learning framework. This approach enriches multimodal embeddings with spatial understanding. We propose two variants: SpatialViLT and MaskedSpatialViLT, focusing on full and masked object regions, respectively. Additionally, SpatialEnsemble combines both approaches, achieving state-of-the-art accuracy. Our models excel in spatial reasoning categories such as directional, topological, and proximity relations, as demonstrated on the challenging Visual Spatial Reasoning (VSR) dataset. This work represents a significant step in enhancing the spatial intelligence of AI systems, crucial for advanced multimodal understanding and real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。