提出分层控制架构,让模仿学习系统在不确定环境下仍可保证安全运行。
Distributionally Robust Imitation Learning: Layered Control Architecture for Certifiable Autonomy
- 构建分层控制框架,融合两种鲁棒方法应对策略误差与外部扰动。
- 通过设计各层输入输出约束,实现整个控制流程的可验证安全性。
- 适合需要高可靠性、可认证自主系统的研发人员参考。
模仿学习(IL)通过专家示范实现自主行为,相比强化学习更高效,但对分布偏移敏感。在基于IL的反馈控制中,存在两类分布偏移:由策略误差引起,以及由外部扰动和模型误差导致的不确定性。此前提出的泰勒级数模仿学习(TaSIL)和 $\mathcal{L}_1$-分布鲁棒自适应控制(\ellonedrac)分别从不同角度解决该问题。本文提出分布鲁棒模仿策略(DRIP)架构,一种分层控制架构(LCA),整合TaSIL与\ellonedrac。通过合理设计各层的输入输出要求,证明可为整个控制链路提供形式化安全证书。该方案打通了学习型模块(如感知)与可验证建模决策之间的路径,为构建全链条可认证自主系统奠定基础。
原文摘要 · Abstract (English)
Imitation learning (IL) enables autonomous behavior by learning from expert demonstrations. While more sample-efficient than comparative alternatives like reinforcement learning, IL is sensitive to compounding errors induced by distribution shifts. There are two significant sources of distribution shifts when using IL-based feedback laws on systems: distribution shifts caused by policy error and distribution shifts due to exogenous disturbances and endogenous model errors due to lack of learning. Our previously developed approaches, Taylor Series Imitation Learning (TaSIL) and $\mathcal{L}_1$ -Distributionally Robust Adaptive Control (\ellonedrac), address the challenge of distribution shifts in complementary ways. While TaSIL offers robustness against policy error-induced distribution shifts, \ellonedrac offers robustness against distribution shifts due to aleatoric and epistemic uncertainties. To enable certifiable IL for learned and/or uncertain dynamical systems, we formulate \textit{Distributionally Robust Imitation Policy (DRIP)} architecture, a Layered Control Architecture (LCA) that integrates TaSIL and~\ellonedrac. By judiciously designing individual layer-centric input and output requirements, we show how we can guarantee certificates for the entire control pipeline. Our solution paves the path for designing fully certifiable autonomy pipelines, by integrating learning-based components, such as perception, with certifiable model-based decision-making through the proposed LCA approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。