通过视觉轨迹投影提升模型对空间时间信息的联合理解能力。
Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding
- 用关键点视觉痕迹投影到深度图,融合时空信息。
- 在SimplerEnv上任务成功率提升4%(比SpatialVLA)和19%(比TraceVLA)。
- 仅需少量训练数据,适合真实场景数据稀缺的应用。
视觉-语言-动作模型在基于视觉观测和文本指令预测智能体在虚拟环境与真实世界中的运动方面展现了显著能力。尽管现有研究多独立增强空间或时间理解,本文提出一种新方法,通过视觉提示整合两者:将观测中关键点的视觉痕迹投影至深度图,使模型能同时捕捉空间与时间信息。实验表明,在SimplerEnv上,该方法使任务成功解决的平均数相比SpatialVLA提升4%,相比TraceVLA提升19%。此外,该增强可在极少训练数据下实现,特别适用于真实世界中数据收集困难的应用场景。
原文摘要 · Abstract (English)
Vision-Language-Action models have demonstrated remarkable capabilities in predicting agent movements within virtual environments and real-world scenarios based on visual observations and textual instructions. Although recent research has focused on enhancing spatial and temporal understanding independently, this paper presents a novel approach that integrates both aspects through visual prompting. We introduce a method that projects visual traces of key points from observations onto depth maps, enabling models to capture both spatial and temporal information simultaneously. The experiments in SimplerEnv show that the mean number of tasks successfully solved increased for 4% compared to SpatialVLA and 19% compared to TraceVLA. Furthermore, we show that this enhancement can be achieved with minimal training data, making it particularly valuable for real-world applications where data collection is challenging. The project page is available at https://ampiromax.github.io/ST-VLA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。