arXiv:2410.22059cs.ROcs.CV2024-10中稿 · WACV2025

用单步生成+视角控制,让机器人零样本重排场景更准更快

PACA: Perspective-Aware Cross-Attention Representation for Zero-Shot Scene Rearrangement

  • 单步融合生成、分割与特征编码,避免多模型误差累积
  • 在真实机器人上实现87%匹配准确率和67%执行成功率
  • 支持6自由度视角控制,突破传统3自由度限制

场景重排(如整理桌面)是机器人操作中的难题,因需预测多样化的物体布局。基于网络规模训练的生成模型(如Stable Diffusion)可通过生成自然场景作为目标来辅助。为便于机器人执行,需提取物级表示以匹配真实场景与生成目标,并计算物体位姿变换。现有方法通常采用多阶段设计,分别使用生成、分割和特征编码模型,易因误差累积导致成功率低;且缺乏对生成目标视角的控制,任务受限于3-自由度设置。本文提出PACA,一种零样本场景重排管道,利用Stable Diffusion生成的视角感知交叉注意力表示。具体而言,我们开发了一种将生成、分割与特征编码整合为单步的表示方法,同时引入视角控制机制,使目标匹配可支持6-自由度相机视图,扩展了以往仅限于3-自由度俯视图的方法。实验表明,该方法在真实机器人上实现零样本性能,在多种场景下平均匹配准确率达87%,执行成功率为67%。

原文摘要 · Abstract (English)

Scene rearrangement, like table tidying, is a challenging task in robotic manipulation due to the complexity of predicting diverse object arrangements. Web-scale trained generative models such as Stable Diffusion can aid by generating natural scenes as goals. To facilitate robot execution, object-level representations must be extracted to match the real scenes with the generated goals and to calculate object pose transformations. Current methods typically use a multi-step design that involves separate models for generation, segmentation, and feature encoding, which can lead to a low success rate due to error accumulation. Furthermore, they lack control over the viewing perspectives of the generated goals, restricting the tasks to 3-DoF settings. In this paper, we propose PACA, a zero-shot pipeline for scene rearrangement that leverages perspective-aware cross-attention representation derived from Stable Diffusion. Specifically, we develop a representation that integrates generation, segmentation, and feature encoding into a single step to produce object-level representations. Additionally, we introduce perspective control, thus enabling the matching of 6-DoF camera views and extending past approaches that were limited to 3-DoF top-down views. The efficacy of our method is demonstrated through its zero-shot performance in real robot experiments across various scenes, achieving an average matching accuracy and execution success rate of 87% and 67%, respectively.

场景重排生成模型6-DoF零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。