arXiv:2509.17783cs.RO2025-09被引 1

让机器人通过交互学习完成复杂操作,实测成功率超79%。

RoboSeek: You Need to Interact with Your Objects

  • 用仿真环境闭环训练,结合视觉先验优化动作策略
  • 在8个长时序任务中平均成功率达79%,远超基准线(<50%)
  • 支持多平台部署,适合需要物理交互的机器人研发

通过探索与交互优化动作执行是机器人操作的潜力方向。然而,针对长时序任务中序列决策、物理约束和感知不确定性等挑战,基于交互的机器人学习方法仍不充分。受具身认知理论启发,我们提出RoboSeek框架,利用交互经验实现具身动作执行。该框架通过仿真中闭环训练优化高层感知模型的先验知识,并借助real2sim2real迁移流程实现真实世界稳健执行。具体而言,我们使用3D重建复现真实环境,构建视觉与物理一致的仿真空间,再基于强化学习与交叉熵方法,结合视觉先验在仿真中训练策略。训练后的策略部署于真实机器人平台执行。RoboSeek具备硬件无关性,在多个机器人平台上评估了8个涉及顺序交互、工具使用与物体操作的长时序任务。实验表明,本方法平均成功率达79%,显著优于基线(均低于50%),验证了其跨任务、跨平台的泛化能力与鲁棒性。结果证实训练框架在复杂动态真实场景中的有效性,以及real2sim2real迁移机制的稳定性,为更通用的具身机器人学习开辟路径。

原文摘要 · Abstract (English)

Optimizing and refining action execution through exploration and interaction is a promising way for robotic manipulation. However, practical approaches to interaction-driven robotic learning are still underexplored, particularly for long-horizon tasks where sequential decision-making, physical constraints, and perceptual uncertainties pose significant challenges. Motivated by embodied cognition theory, we propose RoboSeek, a framework for embodied action execution that leverages interactive experience to accomplish manipulation tasks. RoboSeek optimizes prior knowledge from high-level perception models through closed-loop training in simulation and achieves robust real-world execution via a real2sim2real transfer pipeline. Specifically, we first replicate real-world environments in simulation using 3D reconstruction to provide visually and physically consistent environments, then we train policies in simulation using reinforcement learning and the cross-entropy method leveraging visual priors. The learned policies are subsequently deployed on real robotic platforms for execution. RoboSeek is hardware-agnostic and is evaluated on multiple robotic platforms across eight long-horizon manipulation tasks involving sequential interactions, tool use, and object handling. Our approach achieves an average success rate of 79%, significantly outperforming baselines whose success rates remain below 50%, highlighting its generalization and robustness across tasks and platforms. Experimental results validate the effectiveness of our training framework in complex, dynamic real-world settings and demonstrate the stability of the proposed real2sim2real transfer mechanism, paving the way for more generalizable embodied robotic learning. Project Page: https://russderrick.github.io/Roboseek/

机器人操作具身智能仿真训练长时序任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。