用强化学习自动设计安全的实验信号,提升机电系统参数辨识精度。
Reinforcement Learning for Optimal Experiment Design in Parameter Identification of Mechatronic Systems
- 用强化学习自动生成满足硬件安全约束的激励信号。
- 在10次独立训练中,参数估计精度优于传统方法,安全违规仅0.75%。
- 适合需要自动化、高安全性实验设计的机电系统研究者。
有信息量的激励信号对机电系统精确参数辨识至关重要,但传统系统辨识方法依赖专家知识手工设计信号,且需遵守硬件安全约束,限制了其通用性。我们提出一种强化学习(RL)代理,在Quanser Aero 2测试平台上自主学习最优激励信号,并通过奖励函数设计自动保障安全约束。在10个独立训练种子下评估,该代理在所有三个待辨识参数上均达到具有竞争力的估计精度,优于经典基线方法,且安全违规率仅为0.75%。
原文摘要 · Abstract (English)
Informative excitation signals are critical for accurate system identification of mechatronic systems, yet classical system identification (SI) approaches require expert knowledge and hand-crafted signal design to respect hardware safety constraints, limiting their generalizability. We propose a reinforcement learning (RL) agent that learns optimal excitation signals for a Quanser Aero 2 testbed while autonomously enforcing safety constraints through reward shaping. Evaluated across 10 independent training seeds, our comprehensive agent achieves competitive estimation accuracy across all three identified parameters, outperforming classical baselines while incurring only 0.75% safety violations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。