arXiv:2509.13839cs.RO2025-09

提前预测操作成功概率,避免失败重试,提升机器人操作效率。

Pre-Manipulation Alignment Prediction with Parallel Deep State-Space and Transformer Models

  • 并行使用状态空间模型与Transformer捕捉轨迹时序特征。
  • 在多个数据集上优于现有方法,包括大模型基线。
  • 适合需要高可靠性、低失败率的机器人操作场景。

本文解决开放词汇物体操作任务的未来成功率预测问题。传统方法通常在动作执行后判断成败,难以预防潜在风险,且依赖失败触发重规划,降低操作序列效率。为此,我们提出一种新模型,通过预操作视角图像、计划轨迹与自然语言指令间的对齐程度,提前预测操作成功概率。引入多层级轨迹融合模块,采用先进的深度状态空间模型与Transformer编码器并行处理末端执行器轨迹中的多层级时序自相关性。实验结果表明,该方法在多个基准上均优于现有方法,包括基础模型。

原文摘要 · Abstract (English)

In this work, we address the problem of predicting the future success of open-vocabulary object manipulation tasks. Conventional approaches typically determine success or failure after the action has been carried out. However, they make it difficult to prevent potential hazards and rely on failures to trigger replanning, thereby reducing the efficiency of object manipulation sequences. To overcome these challenges, we propose a model, which predicts the alignment between a pre-manipulation egocentric image with the planned trajectory and a given natural language instruction. We introduce a Multi-Level Trajectory Fusion module, which employs a state-of-the-art deep state-space model and a transformer encoder in parallel to capture multi-level time-series self-correlation within the end effector trajectory. Our experimental results indicate that the proposed method outperformed existing methods, including foundation models.

机器人操作预测模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。