arXiv:2508.19153cs.ROcs.AI2025-08被引 2

用可解释的神经网络提升四足机器人视觉导航能力

QuadKAN: KAN-Enhanced Quadruped Motion Control via End-to-End Reinforcement Learning

  • 用KAN网络构建跨模态控制策略,融合本体感知与视觉信息
  • 在复杂地形上实现更高移动距离、更少碰撞和更低能耗
  • 方法兼具高效性与可解释性,适合机器人控制研究者

我们针对视觉引导的四足机器人运动控制问题,提出基于强化学习的QuadKAN框架。该方法结合本体感知与视觉输入,采用基于样条参数化的柯尔莫戈洛夫-阿诺德网络(KAN)构建跨模态策略。通过样条编码器处理本体感知数据,样条融合头整合多模态输入,其结构与步态的分段光滑特性匹配,显著提升样本效率,降低动作抖动与能耗,并提供可解释的姿态-动作敏感性分析。采用多模态延迟随机化(MMDR)并使用近端策略优化(PPO)进行端到端训练。在多种地形(包括平坦、不平地面及静态/动态障碍物场景)上的评估显示,QuadKAN在收益、移动距离和碰撞率方面均优于现有最先进方法。结果表明,样条参数化策略为鲁棒视觉引导运动提供了简单、有效且可解释的新范式。代码库将在论文接收后公开。

原文摘要 · Abstract (English)

We address vision-guided quadruped motion control with reinforcement learning (RL) and highlight the necessity of combining proprioception with vision for robust control. We propose QuadKAN, a spline-parameterized cross-modal policy instantiated with Kolmogorov-Arnold Networks (KANs). The framework incorporates a spline encoder for proprioception and a spline fusion head for proprioception-vision inputs. This structured function class aligns the state-to-action mapping with the piecewise-smooth nature of gait, improving sample efficiency, reducing action jitter and energy consumption, and providing interpretable posture-action sensitivities. We adopt Multi-Modal Delay Randomization (MMDR) and perform end-to-end training with Proximal Policy Optimization (PPO). Evaluations across diverse terrains, including both even and uneven surfaces and scenarios with static or dynamic obstacles, demonstrate that QuadKAN achieves consistently higher returns, greater distances, and fewer collisions than state-of-the-art (SOTA) baselines. These results show that spline-parameterized policies offer a simple, effective, and interpretable alternative for robust vision-guided locomotion. A repository will be made available upon acceptance.

四足机器人强化学习可解释性多模态控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。