arXiv:2412.19112cs.ROcs.CV2024-12中稿 · CoRL被引 1

基于轨迹和指令提前预测操作成败,提升机器人任务效率

Future Success Prediction in Open-Vocabulary Object Manipulation Tasks Based on End-Effector Trajectories

  • 通过轨迹编码器对操作轨迹加权,捕捉时序动态与交互关系
  • 在RT-1数据集上准确率优于基线方法,实现提前成功预测
  • 适合需要实时决策的开放词汇操作场景,如智能机器人

本研究针对开放词汇物体操作任务中的未来成功预测问题。模型需根据自然语言指令、操作前的视角图像及给定的末端执行器轨迹,预测操作结果。传统方法通常在操作完成后才进行成功率判断,限制了任务整体效率。本文提出一种新方法,通过将轨迹、图像与语言指令对齐,实现操作前的成功预测。引入轨迹编码器,对输入轨迹施加可学习权重,使模型能充分考虑时间动态及物体与末端执行器间的交互关系,从而提升预测准确性。基于RT-1数据集构建大规模基准测试集,实验表明该方法在预测精度上优于现有基线模型。

原文摘要 · Abstract (English)

This study addresses a task designed to predict the future success or failure of open-vocabulary object manipulation. In this task, the model is required to make predictions based on natural language instructions, egocentric view images before manipulation, and the given end-effector trajectories. Conventional methods typically perform success prediction only after the manipulation is executed, limiting their efficiency in executing the entire task sequence. We propose a novel approach that enables the prediction of success or failure by aligning the given trajectories and images with natural language instructions. We introduce Trajectory Encoder to apply learnable weighting to the input trajectories, allowing the model to consider temporal dynamics and interactions between objects and the end effector, improving the model's ability to predict manipulation outcomes accurately. We constructed a dataset based on the RT-1 dataset, a large-scale benchmark for open-vocabulary object manipulation tasks, to evaluate our method. The experimental results show that our method achieved a higher prediction accuracy than baseline approaches.

操作预测轨迹建模多模态机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。