arXiv:2607.11004cs.ROcs.AI2026-07

基于视觉仿生与文本目标的机器人操作规划,支持真实世界泛化。

Affordance-Based Manipulation Planning with Text Goals and Sim-to-Real Generalisation via Real-to-Sim Image Conversion

  • 通过预测动作效果和视觉未来,结合文本目标匹配规划路径。
  • 物体被遮挡时仍能追踪其位置,确保规划连续性。
  • 利用真实图像转模拟图像模块,提升真实场景部署能力。

我们提出一种基于仿生识别与动作效果预测的操作规划系统。该系统以视觉形式推理可能的未来状态,通过多模态目标匹配模块评估候选计划是否符合运行时设定的文本目标。即使物体被遮挡,系统也能通过预测持续追踪文本中指定物体的位置,从而在物体遮挡或初始描述失效的情况下生成有效动作计划。此外,我们引入图像转换模块,将具有不同形状和外观的真实世界物体图像转化为一致视觉表现,以促进物理机器人环境下的操作规划。我们在仿真和硬件上分别评估了各模块性能,并展示了集成系统在一系列挑战性任务中的操作规划能力。

原文摘要 · Abstract (English)

We present a manipulation planning system based on affordance recognition and action effect prediction. The system reasons through possible futures in visual form, and evaluates candidate plans by agreement of predicted outcomes with text-based goals set at run-time, using a multi-modal goal-matching module. Positions of objects named in the goal text are tracked through predictions even when occluded, making it possible to generate action plans even when objects become occluded, or when their initial descriptors cease to identify them in future states. We further expand the system with an image conversion module for translating real-world state images with objects of varied shapes and visual appearances into a consistent visual appearance, to facilitate manipulation planning in a physical robot setup. We evaluate performance of the system's modules in isolation and demonstrate the integrated system's manipulation planning capabilities on a set of challenging tasks in both simulation and on hardware.

操作规划视觉推理真实世界泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。