让机器人用极少示范就学会复杂任务,靠预测未来互动来提升泛化能力。
FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation

- 通过预测任务相关的未来交互特征,实现无需像素级预测的长程推理。
- 在LIBERO上仅用20次示范达95.7%成功率,真实机器人性能提升26%。
- 适合研究少样本机器人学习与视觉-语言-动作模型优化的团队。
视觉-语言-动作(VLA)模型通过大规模多模态预训练实现了通用机器人控制,但在少样本模仿学习下的表现仍有限。我们对主流VLA模型进行了系统性压力测试,发现其性能随示范数据减少而急剧下降,暴露了现有适配策略的关键弱点。为此,我们提出面向未来的条件化框架FOCA,结合任务相关的未来交互嵌入显式预测与未来目标观测的隐式对齐,实现在潜在空间中的长时序推理,无需像素级预测。该方法天然支持与视频世界模型生成的合成视频进行无动作联合训练,并可解释为学习一种未来条件化的价值类表示。大量实验表明,FOCA在LIBERO上使用20个示范即达到95.7%成功率,在RoboCasa上提升7-12%,在真实机器人上实现最高26%的绝对性能增益,建立了少样本VLA适配的新基准。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models enable general-purpose robotic control via large-scale multimodal pretraining, yet their effectiveness under few-shot imitation learning remains limited. We conduct a systematic stress test of state-of-the-art VLA models and show that performance degrades sharply as demonstrations are reduced, revealing a key weakness of existing adaptation strategies. To address this, we introduce FOCA, a future-oriented conditioning framework for data-efficient VLA adaptation. FOCA combines explicit prediction of task-grounded future interaction embeddings with implicit alignment to future goal observations, enabling long-horizon reasoning in latent space without pixel-level prediction. This formulation naturally supports action-free co-training with synthetic videos from video world models and can be interpreted as learning a future-conditioned value-like representation. Extensive experiments demonstrate FOCA achieves 95.7% success with 20 demonstrations on LIBERO, improves 7-12% on RoboCasa, and delivers up to 26% absolute gains on real robots, establishing a new state of the art in few-shot VLA adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。