动态物体操作新框架,实现快速感知与连续控制。
DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation
- 用0.4B小模型+卷积视觉编码器,提升多模态推理速度。
- 连续推理与动作流设计,响应速度更快、适应更及时。
- 自建20万条合成数据集,支持真实世界快速部署。
动态物体操纵仍是视觉-语言-动作(VLA)模型的难题,尽管在静态操作中表现良好,但在需要快速感知、时间预测和持续控制的动态场景中仍面临挑战。我们提出DynamicVLA,通过三项核心设计实现动态物体操纵:1)采用0.4B小型VLA模型搭配卷积视觉编码器,实现空间高效且结构忠实的编码,支持快速多模态推理;2)连续推理机制,使推理与执行重叠,降低延迟并及时适应物体运动;3)基于潜在表示的动作流,通过时间对齐动作执行,弥合感知与执行的鸿沟。为填补动态操作数据空白,我们构建了动态物体操纵(DOM)基准,通过自动化数据采集管道,在2800个场景、206种物体上生成20万条合成数据,并可无遥控快速采集2000条真实世界数据。大量实验表明,该模型在响应速度、感知能力和泛化性能上均有显著提升,成为跨形态通用动态物体操纵的统一框架。
原文摘要 · Abstract (English)
Manipulating dynamic objects remains an open challenge for Vision-Language-Action (VLA) models, which, despite strong generalization in static manipulation, struggle in dynamic scenarios requiring rapid perception, temporal anticipation, and continuous control. We present DynamicVLA, a framework for dynamic object manipulation that integrates temporal reasoning and closed-loop adaptation through three key designs: 1) a compact 0.4B VLA using a convolutional vision encoder for spatially efficient, structurally faithful encoding, enabling fast multimodal inference; 2) Continuous Inference, enabling overlapping reasoning and execution for lower latency and timely adaptation to object motion; and 3) Latent-aware Action Streaming, which bridges the perception-execution gap by enforcing temporally aligned action execution. To fill the missing foundation of dynamic manipulation data, we introduce the Dynamic Object Manipulation (DOM) benchmark, built from scratch with an auto data collection pipeline that efficiently gathers 200K synthetic episodes across 2.8K scenes and 206 objects, and enables fast collection of 2K real-world episodes without teleoperation. Extensive evaluations demonstrate remarkable improvements in response speed, perception, and generalization, positioning DynamicVLA as a unified framework for general dynamic object manipulation across embodiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。