用视觉大模型提升机器人在厨房环境中的物体交互能力
Integrating Vision Foundation Models with Reinforcement Learning for Enhanced Object Interaction
- 将SAM与YOLOv5融合感知,配合PPO强化学习控制
- 交互成功率提升52.5%,导航效率提高33%,奖励增68%
- 适合研究智能机器人感知与决策的开发者
本文提出一种新方法,将视觉基础模型与强化学习结合,以增强模拟环境中对物体的交互能力。通过在AI2-THOR仿真环境中,将分割一切模型(SAM)和YOLOv5与近端策略优化(PPO)智能体结合,使智能体能更有效地感知并操作物体。在四个不同的室内厨房场景中进行的全面实验表明,相比不具备先进感知能力的基线智能体,该方法显著提升了物体交互成功率和导航效率。结果表明,平均累积奖励提升68%,物体交互成功率提高52.5%,导航效率增加33%。这些发现凸显了将基础模型与强化学习结合在复杂机器人任务中的潜力,为构建更强大、更智能的自主代理铺平道路。
原文摘要 · Abstract (English)
This paper presents a novel approach that integrates vision foundation models with reinforcement learning to enhance object interaction capabilities in simulated environments. By combining the Segment Anything Model (SAM) and YOLOv5 with a Proximal Policy Optimization (PPO) agent operating in the AI2-THOR simulation environment, we enable the agent to perceive and interact with objects more effectively. Our comprehensive experiments, conducted across four diverse indoor kitchen settings, demonstrate significant improvements in object interaction success rates and navigation efficiency compared to a baseline agent without advanced perception. The results show a 68% increase in average cumulative reward, a 52.5% improvement in object interaction success rate, and a 33% increase in navigation efficiency. These findings highlight the potential of integrating foundation models with reinforcement learning for complex robotic tasks, paving the way for more sophisticated and capable autonomous agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。