用动态模型让机器人更聪明地完成复杂操作
DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation
- 根据任务需求自动切换机械臂或夹爪的控制模型
- 在模拟和真实环境中成功率显著优于现有方法
- 适合需要精细动作的仓储、医疗等场景
在仓库、医院和家庭等动态环境中,机器人需无缝切换粗略运动与精细操作以完成复杂任务。当前基于预训练视觉语言模型的视觉-语言-动作(VLA)框架难以兼顾任务泛化性与操作精确性。为此,我们提出DAM-VLA,一种基于动态动作模型的VLA框架。该框架将视觉语言模型推理与专用于机械臂和夹爪控制的扩散模型相结合,引入三项创新:(i) 基于任务特定视觉与语言线索的动作路由机制,选择合适动作模型;(ii) 动态动作模型,融合高层语义认知与低层视觉特征进行动作预测;(iii) 双尺度动作加权机制,实现机械臂运动与夹爪操作的动态协调。在广泛评估中,DAM-VLA在模拟环境(SIMPLER、FurnitureBench)和真实场景下均显著优于现有先进VLA方法,在标准抓放任务至长时序、高接触任务中展现强泛化能力。
原文摘要 · Abstract (English)
In dynamic environments such as warehouses, hospitals, and homes, robots must seamlessly transition between gross motion and precise manipulations to complete complex tasks. However, current Vision-Language-Action (VLA) frameworks, largely adapted from pre-trained Vision-Language Models (VLMs), often struggle to reconcile general task adaptability with the specialized precision required for intricate manipulation. To address this challenge, we propose DAM-VLA, a dynamic action model-based VLA framework. DAM-VLA integrates VLM reasoning with diffusion-based action models specialized for arm and gripper control. Specifically, it introduces (i) an action routing mechanism, using task-specific visual and linguistic cues to select appropriate action models (e.g., arm movement or gripper manipulation), (ii) a dynamic action model that fuses high-level VLM cognition with low-level visual features to predict actions, and (iii) a dual-scale action weighting mechanism that enables dynamic coordination between the arm-movement and gripper-manipulation models. Across extensive evaluations, DAM-VLA achieves superior success rates compared to state-of-the-art VLA methods in simulated (SIMPLER, FurnitureBench) and real-world settings, showing robust generalization from standard pick-and-place to demanding long-horizon and contact-rich tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。