多视角感知提升机器人动作进度预测准确率
Multiview Progress Prediction of Robot Activities
- 采用多视角视觉架构捕捉机器人自遮挡下的动作进展
- 在Mobile ALOHA上实现显著优于单视角的预测精度
- 适合需要实时人机协作的智能机器人系统
为使机器人能有效安全地与人类协同工作,必须能够理解正在进行动作的进展。这种能力被称为动作进度预测,对及时协助和自主决策等任务至关重要。然而,机器人动作进展建模常被忽视。此外,单一摄像头难以充分理解机器人的自我动作,因自遮挡会严重影响感知与模型表现。本文提出一种用于机器人操作任务中动作进度预测的多视角架构。在Mobile ALOHA上的实验验证了该方法的有效性。
原文摘要 · Abstract (English)
For robots to operate effectively and safely alongside humans, they must be able to understand the progress of ongoing actions. This ability, known as action progress prediction, is critical for tasks ranging from timely assistance to autonomous decision-making. However, modeling action progression in robotics has often been overlooked. Moreover, a single camera may be insufficient for understanding robot's ego-actions, as self-occlusion can significantly hinder perception and model performance. In this paper, we propose a multi-view architecture for action progress prediction in robot manipulation tasks. Experiments on Mobile ALOHA demonstrate the effectiveness of the proposed approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。