仅凭一次人类示范,就能让机器人学会多种折纸动作。
Instant-Fold: In-Context Imitation Learning for Deformable Object Manipulation

- 用视频对比学习提取变形物体的视觉特征,再结合演示预测动作。
- 在仿真中训练后零样本迁移到真实世界,无需额外数据或微调。
- 适合研究柔性物体操作、具身智能与少样本模仿学习的学者。
柔性物体操作(DOM)因状态高维、部分可观测且交互过程长、拓扑变化频繁而极具挑战。本文提出 Instant-Fold,一种基于上下文模仿学习的框架。仅需一次人类示范,该策略即可直接从示范中推断并执行多种操作模式,包括空间执行方式和顺序的差异,无需梯度更新。方法首先通过时间对比预训练学习形变感知的视觉表征,随后使用基于示范条件的流匹配变换器策略预测动作以实现目标操作模式。整个模型在仿真环境中训练,能泛化至多样折叠模式,并实现零样本迁移至真实场景,无需额外数据采集或微调。视频展示见 https://instant-fold.github.io。
原文摘要 · Abstract (English)
Deformable object manipulation (DOM) is challenging due to high-dimensional, partially observable states that evolve through long-horizon, topology-changing interactions with multiple valid manipulation modes. We introduce Instant-Fold, an in-context imitation learning framework for DOM. Given a single human demonstration, our policy infers and executes diverse manipulation modes directly from the demonstration, including variations in spatial execution and ordering, without requiring gradient updates. Our approach first learns deformation-aware visual representations via temporal contrastive pretraining, after which a flow-matching transformer policy conditioned on the demonstration predicts actions to execute the intended manipulation mode. Trained entirely in simulation, Instant-Fold generalizes across diverse folding modes and transfers zero-shot to real-world settings without additional data collection or finetuning. Videos are available at https://instant-fold.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。