用强化学习设计轨迹,智能分配有限测量预算,提升数字孪生校准精度。
Trajectory Design and Budgeted Querying for Digital Twin Calibration

- 用激励导向的强化学习生成有效轨迹,提升参数估计效率。
- 三次测量预算下终端误差仅0.0092,远优于未校准孪生的0.2031。
- 适用于数据稀缺场景,适合做数字孪生校准研究者参考。
数字孪生校准需要昂贵的交互数据。本文研究两个采集决策:生成哪些轨迹,以及在有限预算下何时进行高价值参数测量。框架结合激励导向的强化学习控制器、带预测不确定性的循环参数估计算器和预算约束的查询策略。在单摆系统中,基于任务的轨迹仅弱恢复重力,无法恢复质量或长度;而使用激励轨迹训练的GRU模型,无查询时平均绝对误差达0.0066。随后在轨迹中途中断全量观测,要求孪生模型依赖估计算法运行剩余部分。估计算法与策略联合使用,在三查询预算下终端误差为0.0092,而未校准孪生误差为0.2031。在部分可观测的Waterworld环境中,五种控制器产生不同隐藏参数的观测误差,混合训练的估计算器实现约4-5%的在线归一化误差。这些探索性案例虽非受控消融实验,但支持将轨迹设计与查询分配作为数据稀缺校准中的显式设计变量。
原文摘要 · Abstract (English)
Digital-twin calibration requires interaction data that is expensive to collect. We study two acquisition decisions: which trajectories to generate, and when to spend a limited budget on privileged parameter measurements. Our framework couples an excitation-oriented reinforcement learning controller, a recurrent parameter estimator with predictive uncertainty, and a budgeted query policy. In Pendulum, a Random Forest diagnostic recovers gravity only weakly from task-oriented trajectories and does not recover mass or length, while a GRU trained on excitation-oriented trajectories reaches a mean absolute error of 0.0066 with no queries. We then withdraw continuous oracle access partway through an episode, so that the twin must run on the estimator's output for the remainder. The estimator-plus-policy pipeline achieves a terminal error of 0.0092 under a three-query budget, against 0.2031 for an uncalibrated twin. In partially observable Waterworld, five controllers produce different observed error profiles across three hidden parameters, and an estimator trained on a five-controller mixture reaches online normalized errors of roughly 4-5%. These exploratory case studies are not controlled ablations, but they motivate treating trajectory design and query allocation as explicit design variables in data-scarce calibration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。