通过信念空间规划提升模型识别精度,实现更安全的自适应控制。
Model Identification Adaptive Control with $ρ$-POMDP Planning
- 将模型识别与控制结合为信念空间规划问题,用ρ-POMDP建模参数不确定性
- 在小车摆杆和飞机飞行任务中,比传统方法更抗参数突变干扰
- 适合需要高可靠性模型识别的机器人与航空航天控制系统
精确的系统建模对安全有效的控制至关重要,错误识别会导致误差累积,尤其在部分可观测条件下。本文将信息性输入设计与模型识别自适应控制(MIAC)建模为信念空间规划问题,采用带信念依赖奖励的局部可观测马尔可夫决策过程(ρ-POMDP)。将系统参数视为需定位的隐藏状态,同时实现控制与参数估计。使用改进的信念空间迭代线性二次调节器(BiLQR)求解。在完全可观测和部分可观测的小车摆杆与稳定飞行飞机任务中验证了该方法的有效性。相比回归、滤波和局部最优控制基线方法,本方法在系统参数瞬时扰动下仍表现更优。
原文摘要 · Abstract (English)
Accurate system modeling is crucial for safe, effective control, as misidentification can lead to accumulated errors, especially under partial observability. We address this problem by formulating informative input design and model identification adaptive control (MIAC) as belief space planning problems, modeled as partially observable Markov decision processes with belief-dependent rewards ($ρ$-POMDPs). We treat system parameters as hidden state variables that must be localized while simultaneously controlling the system. We solve this problem with an adapted belief-space iterative Linear Quadratic Regulator (BiLQR). We demonstrate it on fully and partially observable tasks for cart-pole and steady aircraft flight domains. Our method outperforms baselines such as regression, filtering, and local optimal control methods, even under instantaneous disturbances to system parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。