用隐马尔可夫模型解决机器人操作中的状态混淆问题
SAGE:State-Aware Guided End-to-End Policy for Multi-Stage Sequential Tasks via Hidden Markov Decision Process
- 将任务建模为隐马尔可夫决策过程,通过隐状态消除视觉相似下的动作歧义
- 真实场景中多任务测试成功率达100%,远超基线方法
- 仅需约13%状态的人工标注,适合低标注成本的复杂任务学习
多阶段序列机器人操作任务在机器人领域普遍存在且至关重要。这类任务常面临状态模糊问题,即视觉上相似的观察对应不同动作。本文提出SAGE,一种基于隐马尔可夫决策过程(HMDP)的状态感知引导模仿学习框架,显式建模任务的潜在阶段以解决模糊性。通过状态转移网络推断隐藏状态,并设计状态感知动作策略,该策略同时依赖观测与隐藏状态生成动作,从而实现跨阶段的歧义消解。为减少人工标注负担,提出结合主动学习与软标签插值的半自动标注流程。在多个具有状态模糊性的复杂真实任务中,SAGE在标准评估协议下实现100%任务成功率,显著优于基线方法。消融实验表明,仅需约13%状态的手动标注即可保持优异性能,证明其高效性。
原文摘要 · Abstract (English)
Multi-stage sequential (MSS) robotic manipulation tasks are prevalent and crucial in robotics. They often involve state ambiguity, where visually similar observations correspond to different actions. We present SAGE, a state-aware guided imitation learning framework that models tasks as a Hidden Markov Decision Process (HMDP) to explicitly capture latent task stages and resolve ambiguity. We instantiate the HMDP with a state transition network that infers hidden states, and a state-aware action policy that conditions on both observations and hidden states to produce actions, thereby enabling disambiguation across task stages. To reduce manual annotation effort, we propose a semi-automatic labeling pipeline combining active learning and soft label interpolation. In real-world experiments across multiple complex MSS tasks with state ambiguity, SAGE achieved 100% task success under the standard evaluation protocol, markedly surpassing the baselines. Ablation studies further show that such performance can be maintained with manual labeling for only about 13% of the states, indicating its strong effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。