arXiv:2604.11674cs.ROcs.AI2026-04

用自然语言生成任务相关抓取数据,提升机器人操作的泛化能力

AffordSim: A Scalable Data Generator and Benchmark for Affordance-Aware Robotic Manipulation

论文配图:AffordSim: A Scalable Data Generator and Benchmark for Affordance-Aware Robotic Manipulation
图 1 · 摘自论文原文
  • 基于视觉-语言-动作模型,自动推理物体功能区域并生成合理抓取点
  • 在50个任务上达到人工标注93%的轨迹成功率,真实机器人零样本迁移成功率达24%
  • 支持多种机器人和物体,适用于需要理解物体功能的复杂操作场景

许多日常机器人操作技能依赖于物体的功能属性,其成功与否取决于是否接触后续动作所需的特定功能区域。现有仿真数据生成方法要么使用通用抓取估计算法,忽略任务语义导致抓取位置不匹配,要么依赖手动标注,难以扩展。为此,我们提出AffordSim,一个集成开放词汇3D功能属性预测的可扩展数据生成器与基准测试平台。给定自然语言任务描述,AffordSim合成任务相关的场景,发出功能属性查询,将查询定位到物体表面,采样区域条件下的抓取姿态,并通过运动规划筛选可执行方案。同时,随机化物体位姿、纹理、光照、图像噪声及跨视角背景,以增强从仿真到现实的迁移能力。我们构建了包含50个任务、5种机器人形态和500多个刚体与铰接物体的基准测试集。AffordSim在功能关键任务上达到人工标注93%的轨迹收集成功率,在复杂复合任务上为89%。基于AffordSim数据训练的视觉-语言-动作策略在真实Franka FR3机器人上实现24%的平均成功率,且无需微调。

原文摘要 · Abstract (English)

Many everyday robot manipulation skills are affordance-dependent, with success determined by whether the robot contacts the functional object region required by the subsequent action. Current simulation data generators obtain contacts from generic grasp estimators or per-object manual contact annotations, but generic estimators rank stable grasps without task semantics and often select contacts that are misaligned with the downstream action, while manual contact annotations must be rewritten for each new object and task. To solve these challenges, we introduce AffordSim, a scalable data generator and benchmark that integrates open-vocabulary 3D affordance prediction into simulation-based trajectory generation. Given a natural-language task description, AffordSim synthesizes a task-relevant scene, emits affordance queries, grounds them on object surfaces, samples region-conditioned grasps, and selects executable candidates with motion planning. It further randomizes object pose, texture, lighting, image noise, and cross-viewpoint backgrounds for sim-to-real transfer. We instantiate AffordSim as a 50-task benchmark across diverse manipulation skills, five robot embodiments, and 500+ rigid and articulated objects. AffordSim achieves 93% of the trajectory collection success rate of manual contact annotations on affordance-critical tasks and 89% on hard composite tasks. Vision-language-action policies trained on AffordSim data transfer zero-shot to a real Franka FR3, reaching 24% average success.

机器人操作功能感知仿真生成零样本迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。