用视觉大模型+记忆架构,让机器人在复杂环境中更智能地完成多任务操作。
SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation
- 基于多视角视觉特征和记忆机制的机器人策略网络
- 在RLBench上达86.8%成功率,内存任务上达94.3%
- 新基准MemoryBench可评估机器人空间记忆能力
在多样化、动态环境中运行的机器人操作系必须具备多任务交互、泛化到未见场景以及空间记忆能力。尽管机器人操作已取得显著进展,但现有方法在应对复杂环境变化和记忆依赖性任务时仍显不足。为此,我们提出SAM2Act,一种基于多视图视觉基础模型的机器人变换器策略,采用多分辨率上采样与大规模基础模型的视觉表征。SAM2Act在RLBench基准上实现18个任务平均86.8%的成功率,并在The Colosseum基准上表现出色,环境扰动下性能仅下降4.3%。在此基础上,我们进一步提出SAM2Act+,一种受SAM2启发的记忆架构,包含记忆库、编码器和注意力机制,以增强空间记忆能力。为评估记忆依赖任务,我们引入MemoryBench,一个专门用于评估机器人空间记忆与动作回溯能力的新基准。SAM2Act+在MemoryBench的记忆任务中达到94.3%平均成功率,显著优于现有方法,推动了记忆型机器人系统的发展。
原文摘要 · Abstract (English)
Robotic manipulation systems operating in diverse, dynamic environments must exhibit three critical abilities: multitask interaction, generalization to unseen scenarios, and spatial memory. While significant progress has been made in robotic manipulation, existing approaches often fall short in generalization to complex environmental variations and addressing memory-dependent tasks. To bridge this gap, we introduce SAM2Act, a multi-view robotic transformer-based policy that leverages multi-resolution upsampling with visual representations from large-scale foundation model. SAM2Act achieves a state-of-the-art average success rate of 86.8% across 18 tasks in the RLBench benchmark, and demonstrates robust generalization on The Colosseum benchmark, with only a 4.3% performance gap under diverse environmental perturbations. Building on this foundation, we propose SAM2Act+, a memory-based architecture inspired by SAM2, which incorporates a memory bank, an encoder, and an attention mechanism to enhance spatial memory. To address the need for evaluating memory-dependent tasks, we introduce MemoryBench, a novel benchmark designed to assess spatial memory and action recall in robotic manipulation. SAM2Act+ achieves an average success rate of 94.3% on memory-based tasks in MemoryBench, significantly outperforming existing approaches and pushing the boundaries of memory-based robotic systems. Project page: sam2act.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。