arXiv:2604.10579cs.ROcs.AI2026-04

用3D生成模型创造多样操作数据,让机器人学会泛化抓取新物体。

AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Afford Correspondence

论文配图:AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Afford Correspondence
图 1 · 摘自论文原文
  • 基于3D网格语义关键点对应生成新操作轨迹
  • 真实世界与仿真中对未见物体实现零样本泛化
  • 适合需要高效学习新物体操作的机器人研究者

尽管现代模仿学习在机器人操作中取得进展,其性能仍受限于几何变化带来的数据多样性不足。AffordGen框架利用强大的3D生成模型和视觉基础模型(VFMs),通过大规模3D网格间有意义的关键点语义对应关系,生成新的机器人操作轨迹。该大规模、具备功能感知的数据集用于训练鲁棒的闭环视觉-运动策略,结合了功能的语义泛化能力与端到端学习的响应鲁棒性。仿真与真实世界实验表明,使用AffordGen训练的策略在未见物体上实现高成功率,并具备零样本泛化能力,显著提升机器人学习的数据效率。

原文摘要 · Abstract (English)

Despite the recent success of modern imitation learning methods in robot manipulation, their performance is often constrained by geometric variations due to limited data diversity. Leveraging powerful 3D generative models and vision foundation models (VFMs), the proposed AffordGen framework overcomes this limitation by utilizing the semantic correspondence of meaningful keypoints across large-scale 3D meshes to generate new robot manipulation trajectories. This large-scale, affordance-aware dataset is then used to train a robust, closed-loop visuomotor policy, combining the semantic generalizability of affordances with the reactive robustness of end-to-end learning. Experiments in simulation and the real world show that policies trained with AffordGen achieve high success rates and enable zero-shot generalization to truly unseen objects, significantly improving data efficiency in robot learning. Project Page: https://jiaweiz9.github.io/AffordGen-release/

机器人操作生成模型泛化能力模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。