让机器人听懂动作细节,精准执行复杂指令。
KineVLA: Towards Kinematics-Aware Vision-Language-Action Models with Bi-Level Action Decomposition
- 分两层解耦目标与运动参数,显式建模动作细节
- 在真实机器人上实现更精确、可泛化的操作表现
- 适合需要精细动作控制的智能机器人场景
本文提出一种新型的富含运动学信息的视觉-语言-动作(VLA)任务,其中语言指令在关键节点密集编码方向、轨迹、姿态和相对位移等多样运动学属性,而任务目标保持不变,执行轨迹需随指令中的运动学要求动态调整。为应对这一挑战,我们提出KineVLA框架,通过双层动作表示与双层推理令牌,显式分离目标不变性与运动学可变性,作为语言与动作对齐的监督中间变量。为此,我们构建了涵盖仿真与真实机器人平台的运动学感知VLA数据集,包含指令级运动学变化与双层标注。在LIBERO及Realman-75机器人上的大量实验表明,KineVLA在运动学敏感基准上持续优于强基线,实现更精确、可控且泛化能力更强的操作行为。
原文摘要 · Abstract (English)
In this paper, we introduce a novel kinematics-rich vision-language-action (VLA) task, in which language commands densely encode diverse kinematic attributes (such as direction, trajectory, orientation, and relative displacement) from initiation through completion, at key moments, unlike existing action instructions that capture kinematics only coarsely or partially, thereby supporting fine-grained and personalized manipulation. In this setting, where task goals remain invariant while execution trajectories must adapt to instruction-level kinematic specifications. To address this challenge, we propose KineVLA, a vision-language-action framework that explicitly decouples goal-level invariance from kinematics-level variability through a bi-level action representation and bi-level reasoning tokens to serve as explicit, supervised intermediate variables that align language and action. To support this task, we construct the kinematics-aware VLA datasets spanning both simulation and real-world robotic platforms, featuring instruction-level kinematic variations and bi-level annotations. Extensive experiments on LIBERO and a Realman-75 robot demonstrate that KineVLA consistently outperforms strong VLA baselines on kinematics-sensitive benchmarks, achieving more precise, controllable, and generalizable manipulation behaviors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。