用物理约束生成数据,解决太空任务中样本稀缺问题
Mitigating Data Scarcity in Spaceflight Applications for Offline Reinforcement Learning Using Physics-Informed Deep Generative Models

- 构建物理感知的变分自编码器,学习真实轨迹与物理模型差异
- 生成数据使强化学习策略成功率提升,优于传统方法
- 适合航天器控制等高成本、低数据场景的智能系统设计
在物理系统部署强化学习控制器常受仿真到现实(sim-to-real)差距限制,尤其在太空飞行中,因成本高昂和行星探测数据稀缺而难以获取真实训练数据。传统方法如系统辨识和合成数据生成依赖充足数据,且常因建模假设或缺乏物理约束而失效。本文提出在生成模型中引入物理学习偏置,开发基于互信息的分离变分自编码器(MI-VAE),学习观测轨迹与物理模型预测轨迹之间的差异。其潜在空间可生成符合物理规律的合成数据。在有限真实数据下的行星着陆任务中评估显示,使用MI-VAE生成的数据增强后,下游强化学习性能显著提升,在统计保真度、样本多样性及策略成功率上均优于标准VAE。该工作为复杂、数据受限环境中的自主控制器鲁棒性提升提供了可扩展方案。
原文摘要 · Abstract (English)
The deployment of reinforcement learning (RL)-based controllers on physical systems is often limited by poor generalization to real-world scenarios, known as the simulation-to-reality (sim-to-real) gap. This gap is particularly challenging in spaceflight, where real-world training data are scarce due to high cost and limited planetary exploration data. Traditional approaches, such as system identification and synthetic data generation, depend on sufficient data and often fail due to modeling assumptions or lack of physics-based constraints. We propose addressing this data scarcity by introducing physics-based learning bias in a generative model. Specifically, we develop the Mutual Information-based Split Variational Autoencoder (MI-VAE), a physics-informed VAE that learns differences between observed system trajectories and those predicted by physics-based models. The latent space of the MI-VAE enables generation of synthetic datasets that respect physical constraints. We evaluate MI-VAE on a planetary lander problem, focusing on limited real-world data and offline RL training. Results show that augmenting datasets with MI-VAE samples significantly improves downstream RL performance, outperforming standard VAEs in statistical fidelity, sample diversity, and policy success rate. This work demonstrates a scalable strategy for enhancing autonomous controller robustness in complex, data-constrained environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。