arXiv:2410.15250cs.LG2024-10被引 1

用物理方程增强稀疏观测,让机器人在复杂流体中更准更快地控制。

Multi-modal Policies with Physics-informed Representations in Complex Fluid Environments

  • 融合流体方程与稀疏观测,学习统一状态表示。
  • 即使缺失部分传感器数据,仍保持与真实状态高度一致。
  • 适合水下机器人、航天器等复杂流体环境的控制任务。

流体环境中的控制是众多领域的重要研究方向,如水下机器人、航空航天和生物医学系统。然而,由于传感器限制或故障,实际观测常出现稀疏或缺失问题,且模态不一(如速度与压力传感器)。本文提出物理信息表示(PIR)算法,用于多模态控制策略,以利用复杂流体环境中的稀疏随机观测。PIR将稀疏观测数据与偏微分方程(PDE)信息结合,提炼流体系统的统一表示。核心思想是:PDE解由方程、初值和边界条件决定;给定方程后,仅需学习初值与边界条件的表示,即可定义特定流体系统的演化轨迹。具体通过PDE损失拟合神经网络,并结合带有随机数量与多模态观测的数据损失,将初始与边界条件信息传播至表示中。表示作为可学习参数或编码器输出。实验表明,即使缺失部分模态,PIR仍显著优于基线方法,与真实特征保持更高一致性。此外,将PIR与强化学习结合,在控制任务中使机器人从随机起点快速准确穿越复杂涡街,抵达随机目标。

原文摘要 · Abstract (English)

Control in fluid environments is an important research area with numerous applications across various domains, including underwater robotics, aerospace engineering, and biomedical systems. However, in practice, control methods often face challenges due to sparse or missing observations, stemming from sensor limitations and faults. These issues result in observations that are not only sparse but also inconsistent in their number and modalities (e.g., velocity and pressure sensors). In this work, we propose a Physics-Informed Representation (PIR) algorithm for multi-modal policies of control to leverage the sparse and random observations in complex fluid environments. PIR integrates sparse observational data with the Partial Differential Equation (PDE) information to distill a unified representation of fluid systems. The main idea is that PDE solutions are determined by three elements: the equation, initial conditions, and boundary conditions. Given the equation, we only need to learn the representation of the initial and boundary conditions, which define a trajectory of a specific fluid system. Specifically, it leverages PDE loss to fit the neural network and data loss calculated on the observations with random quantities and multi-modalities to propagate the information with initial and boundary conditions into the representations. The representations are the learnable parameters or the output of the encoder. In the experiments, the PIR illustrates the superior consistency with the features of the ground truth compared with baselines, even when there are missing modalities. Furthermore, PIR combined with Reinforcement Learning has been successfully applied in control tasks where the robot leverages the learned state by PIR faster and more accurately, passing through the complex vortex street from a random starting location to reach a random target.

流体控制物理信息多模态强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。