轻量级模型解决工业场景空间推理难题
TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints
- 分阶段训练+专家混合模块融合多模态特征
- 64M参数模型在工业挑战赛中获第5名,得分66.8861
- 适合资源受限环境下复杂空间理解任务
在仓库级环境中的细粒度空间关系推理对现有视觉语言模型仍具挑战,尤其在理解三维布局、物体排列及多模态线索方面表现有限。本文提出TinyGiantVLM,一种轻量级模块化双阶段框架,专为物理空间推理设计,区别于传统地理推理。该方法利用预训练视觉主干从RGB和深度模态中编码全局与区域级特征。为应对高模态输入复杂性和多样化问题类型,引入混合专家(MoE)融合模块,动态整合空间表征以支持下游推理并提升收敛性。训练采用两阶段策略:第一阶段生成自由形式答案以增强空间推理能力;第二阶段使用标准化答案进行评估。在AI City Challenge 2025 Track 3上的评测显示,64M参数基础模型取得66.8861分,位列第5,展现出在工业环境中连接视觉感知与空间理解的强性能。此外,扩展MoE容量的80M参数变体在空间推理任务中表现更优。
原文摘要 · Abstract (English)
Reasoning about fine-grained spatial relationships in warehouse-scale environments poses a significant challenge for existing vision-language models (VLMs), which often struggle to comprehend 3D layouts, object arrangements, and multimodal cues in real-world industrial settings. In this paper, we present TinyGiantVLM, a lightweight and modular two-stage framework designed for physical spatial reasoning, distinguishing itself from traditional geographic reasoning in complex logistics scenes. Our approach encodes both global and region-level features from RGB and depth modalities using pretrained visual backbones. To effectively handle the complexity of high-modality inputs and diverse question types, we incorporate a Mixture-of-Experts (MoE) fusion module, which dynamically combines spatial representations to support downstream reasoning tasks and improve convergence. Training is conducted in a two-phase strategy: the first phase focuses on generating free-form answers to enhance spatial reasoning ability, while the second phase uses normalized answers for evaluation. Evaluated on Track 3 of the AI City Challenge 2025, our 64M-parameter base model achieved 5th place on the leaderboard with a score of 66.8861, demonstrating strong performance in bridging visual perception and spatial understanding in industrial environments. We further present an 80M-parameter variant with expanded MoE capacity, which demonstrates improved performance on spatial reasoning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。