提升机器人学习策略部署时的可靠性,解决实际应用中的失效问题。
Deployment-Time Reliability of Learned Robot Policies
- 通过运行时监控检测策略行为异常和任务进展偏离,无需故障数据。
- 用影响函数追溯成功与失败对应训练样本,实现策略可解释性诊断。
- 结合成功率估计与可行性规划,保障长序列任务稳定执行。
基于学习的机器人操作虽已取得显著进展,但部署阶段的可靠性仍是现实应用的核心障碍,分布偏移、误差累积及复杂任务依赖共同削弱系统性能。本文研究如何在部署阶段通过围绕策略的机制提升可靠性,提出三类互补方法:第一,开发运行时监控技术,通过识别闭环策略行为不一致和任务进展偏差来预测潜在失败,无需故障数据或任务特定监督;第二,提出以数据为中心的策略可解释框架,利用影响函数将部署时的成功与失败追溯至关键训练示范,实现有原则的诊断与数据集优化;第三,针对长时序任务,将策略协调建模为行为序列成功率的估计与最大化问题,并通过可行性感知的任务规划扩展至开放式语言指令任务。这些工作聚焦部署核心挑战,推动学习型机器人策略在真实场景中可靠、可扩展应用的基础发展,未来持续进步对实现可信自主机器人至关重要。
原文摘要 · Abstract (English)
Recent advances in learning-based robot manipulation have produced policies with remarkable capabilities. Yet, reliability at deployment remains a fundamental barrier to real-world use, where distribution shift, compounding errors, and complex task dependencies collectively undermine system performance. This dissertation investigates how the reliability of learned robot policies can be improved at deployment time through mechanisms that operate around them. We develop three complementary classes of deployment-time mechanisms. First, we introduce runtime monitoring methods that detect impending failures by identifying inconsistencies in closed-loop policy behavior and deviations in task progress, without requiring failure data or task-specific supervision. Second, we propose a data-centric framework for policy interpretability that traces deployment-time successes and failures to influential training demonstrations using influence functions, enabling principled diagnosis and dataset curation. Third, we address reliable long-horizon task execution by formulating policy coordination as the problem of estimating and maximizing the success probability of behavior sequences, and we extend this formulation to open-ended, language-specified tasks through feasibility-aware task planning. By centering on core challenges of deployment, these contributions advance practical foundations for the reliable, real-world use of learned robot policies. Continued progress on these foundations will be essential for enabling trustworthy and scalable robot autonomy in the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。