arXiv:2506.22365cs.LGcs.RO2025-06被引 4

用可读符号程序引导强化学习,实现零样本室内导航

Reinforcement Learning with Physics-Informed Symbolic Program Priors for Zero-Shot Wireless Indoor Navigation

  • 用领域语言表达物理先验,通过符号程序指导神经控制器
  • 相比纯符号或纯神经方法,训练时间减少26%以上
  • 适合需高效、可解释控制策略的智能系统开发者

在强化学习(RL)应对物理控制任务时,融入物理先验的归纳偏置可提升训练样本效率并增强泛化能力。然而,当前方法需大量人工干预和领域知识,难以推广。本文提出一种符号化方法,将物理先验以人类可读的领域特定语言(DSL)形式提炼为归纳偏置。由于导航任务存在观测不完整与物理约束,这些DSL先验无法直接转化为可执行策略。为此,我们构建了物理信息程序引导的强化学习框架(PiPRL),采用分层模块化神经符号集成:元符号程序接收神经感知模块提取的语义特征,形成符号编程基础,编码物理先验并引导低层神经控制器的强化学习过程。大量实验表明,PiPRL持续优于纯符号或纯神经策略,在程序引导下训练时间减少超过26%。

原文摘要 · Abstract (English)

When using reinforcement learning (RL) to tackle physical control tasks, inductive biases that encode physics priors can help improve sample efficiency during training and enhance generalization in testing. However, the current practice of incorporating these helpful physics-informed inductive biases inevitably runs into significant manual labor and domain expertise, making them prohibitive for general users. This work explores a symbolic approach to distill physics-informed inductive biases into RL agents, where the physics priors are expressed in a domain-specific language (DSL) that is human-readable and naturally explainable. Yet, the DSL priors do not translate directly into an implementable policy due to partial and noisy observations and additional physical constraints in navigation tasks. To address this gap, we develop a physics-informed program-guided RL (PiPRL) framework with applications to indoor navigation. PiPRL adopts a hierarchical and modularized neuro-symbolic integration, where a meta symbolic program receives semantically meaningful features from a neural perception module, which form the bases for symbolic programming that encodes physics priors and guides the RL process of a low-level neural controller. Extensive experiments demonstrate that PiPRL consistently outperforms purely symbolic or neural policies and reduces training time by over 26% with the help of the program-based inductive biases.

强化学习符号学习室内导航神经符号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。