用稀疏轨迹避免机器人模型遗忘,零样本泛化更强。
NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
- 聚焦机械臂轨迹,通过时空压缩减少训练数据密度。
- 计算量低于pi0一个数量级,零样本性能更优。
- 无需手腕摄像头,适配多平台且保留语言理解能力。
视觉-语言-动作(VLA)模型在具身智能中取得重要进展,但面临灾难性遗忘等现实部署障碍,主要源于对连续动作序列的过度依赖,导致知识隔离。为此,我们提出窄化轨迹VLA(NoTVLA)框架:将关注点聚焦于稀疏轨迹,避免密集轨迹微调带来的遗忘问题。核心创新在于轨迹规划策略——不以目标物体轨迹为中心,而是针对机械臂末端执行器轨迹进行时间压缩与空间推理剪枝。训练采用稀疏轨迹而非密集动作轨迹,显著提升零样本性能。在多任务评估中,NoTVLA表现优于pi0,且计算资源消耗低于pi0一个数量级,无需腕部摄像头。该设计使操作精度接近单任务专家模型,同时保留语言能力,支持跨平台统一部署,并可在新视角下实现零样本泛化。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models represent a pivotal advance in embodied intelligence, yet they confront critical barriers to real-world deployment, most notably catastrophic forgetting. This issue stems from their overreliance on continuous action sequences or action chunks, which inadvertently create isolated data silos that disrupt knowledge retention across tasks. To tackle these challenges, we propose the Narrowing of Trajectory VLA (NoTVLA) framework: a novel approach that narrows its focus to sparse trajectories, thereby avoiding the catastrophic forgetting associated with dense trajectory fine-tuning. A key innovation of NoTVLA lies in its trajectory planning strategy: instead of centering on the target object's trajectory, it leverages temporal compression and spatial reasoning pruning specifically for the robot end effector's trajectory. Furthermore, training is conducted using these sparse trajectories rather than dense action trajectories, an optimization that delivers remarkable practical advantages with better performance in zero-shot. In multi-task evaluation scenarios, NoTVLA achieves superior performance and generalization compared to pi0 while operating under two critical constraints: it uses over an order of magnitude less computing power than pi0 and requires no wrist-mounted camera. This design ensures that NoTVLA's operational accuracy closely approximates that of single-task expert models. Crucially, it also preserves the model's inherent language capabilities, enabling zero-shot generalization in specific scenarios, supporting unified model deployment across multiple robot platforms, and fostering a degree of generalization even when perceiving tasks from novel perspectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。