用3D模拟训练大模型动态空间推理能力,效果超越真实图像数据。
SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
- 构建17.5万条问答的动态空间模拟数据集,涵盖运动与位置关系变化。
- 在真实图像测试集上,模型准确率提升11%,长视频推理也显著改善。
- 模拟标注比伪标注更有效,适合视觉语言模型的空间认知训练。
运动与空间推理是多个现实应用所需的基础认知能力。尽管已有研究指出大模型在静态空间关系推理上表现不佳,但多数未涉及动态运动感知,即对自身运动与物体运动如何影响空间关系的推理。手动标注物体和相机运动成本高昂。为此,我们提出SAT,一个基于3D模拟器构建的动态空间素养训练数据集,包含17.5万对问答和2万场景,覆盖静态与动态空间推理。同时,我们还构建了一个小规模(150个图像问答)但极具挑战性的真实图像动态空间测试集。结合6个现有静态空间基准,我们系统评估了提升静态与动态空间意识的方法。结果表明,模拟训练能有效提升模型空间能力并迁移到真实图像。完美标注的模拟数据比伪标注真实图像更有效。例如,SAT训练使LLaVA-13B模型在多个基准上平均提升11%,使LLaVA-Video-7B模型平均提升8%,甚至优于部分大型专有模型。虽然静态推理随合成数据训练有所改进,但动态推理仍有巨大提升空间。
原文摘要 · Abstract (English)
Reasoning about motion and space is a fundamental cognitive capability that is required by multiple real-world applications. While many studies highlight that large multimodal language models (MLMs) struggle to reason about space, they only focus on static spatial relationships, and not dynamic awareness of motion and space, i.e., reasoning about the effect of egocentric and object motions on spatial relationships. Manually annotating such object and camera movements is expensive. Hence, we introduce SAT, a simulated spatial aptitude training dataset utilizing 3D simulators, comprising both static and dynamic spatial reasoning across 175K question-answer (QA) pairs and 20K scenes. Complementing this, we also construct a small (150 image-QAs) yet challenging dynamic spatial test set using real-world images. Leveraging our SAT datasets and 6 existing static spatial benchmarks, we systematically investigate what improves both static and dynamic spatial awareness. Our results reveal that simulations are surprisingly effective at imparting spatial aptitude to MLMs that translate to real images. We show that perfect annotations in simulation are more effective than existing approaches of pseudo-annotating real images. For instance, SAT training improves a LLaVA-13B model by an average 11% and a LLaVA-Video-7B model by an average 8% on multiple spatial benchmarks, including our real-image dynamic test set and spatial reasoning on long videos -- even outperforming some large proprietary models. While reasoning over static relationships improves with synthetic training data, there is still considerable room for improvement for dynamic reasoning questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。