arXiv:2603.12807cs.ROcs.SY2026-03

用强化学习控制受力受限的椭圆柱体,验证了复杂运动可行性

Reinforcement Learning for Elliptical Cylinder Motion Control Tasks

  • 基于强化学习求解椭圆柱体在有限扭矩下的运动控制问题
  • 高质心或长宽比大的椭圆柱难以实现竖直稳定或半圈旋转
  • 对比经典两阶段控制器,展现学习方法在复杂场景的潜力

由于输入受限带来的挑战,设备的控制始终是研究热点。例如倒立摆是控制理论和机器学习中的基准问题。本文聚焦于椭圆柱体在有限扭矩下的运动控制,其灵感来源于远程磁驱动设备,因距离限制只能施加有限扭矩。本文目标是定义并解决椭圆柱体在受限输入扭矩下的控制问题,采用强化学习方法。作为经典基线,评估了一个由能量整形摆起控制律与局部线性二次调节器(LQR)稳定器组成的两阶段控制器。摆起控制器通过提升系统机械能将状态引导至目标平衡点邻域,对非线性模型进行线性化后得到一个有界输入的LQR,用于调节角度和角速度至目标姿态。该摆起+LQR策略是欠驱动系统中强而可解释的参考,可用于与学习策略在相同约束和参数下的比较。结果表明,虽然学习可行,但在质量增大或长宽比强烈不等的情况下,实现竖直稳定或半圈旋转仍非常困难。

原文摘要 · Abstract (English)

The control of devices with limited input always bring attention to solve by research due to its difficulty and non-trival solution. For instance, the inverted pendulum is benchmarking problem in control theory and machine learning. In this work, we are focused on the elliptical cylinder and its motion under limited torque. The inspiration of the problem is from untethered magnetic devices, which due to distance have to operate with limited input torque. In this work, the main goal is to define the control problem of elliptic cylinder with limited input torque and solve it by Reinforcement Learning. As a classical baseline, we evaluate a two-stage controller composed of an energy-shaping swing-up law and a local Linear Quadratic Regulator (LQR) stabilizer around the target equilibrium. The swing-up controller increases the system's mechanical energy to drive the state toward a neighborhood of the desired equilibrium, a linearization of the nonlinear model yields an LQR that regulates the angle and angular-rate states to the target orientation with bounded input. This swing-up + LQR policy is a strong, interpretable reference for underactuated system and serves a point of comparison to the learned policy under identical limits and parameters. The solution shows that the learning is possible however, the different cases like stabilization in upward position or rotating of half turn are very difficult for increasing mass or ellipses with a strongly unequal perimeter ratio.

强化学习控制理论欠驱动系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。