arXiv:2608.17596cs.ROcs.AI2026-08

小体积机器人用最少预设知识,自主学会复杂运动技能。

tinyDSM: A Framework for Skill Modeling and Development for Resource-Constrained Millirobots

论文配图:tinyDSM: A Framework for Skill Modeling and Development for Resource-Constrained Millirobots
图 1 · 摘自论文原文
  • 用内在动机驱动学习,结合知识图谱与运动推理实现渐进式技能演化。
  • 36cm³的微型机器人在15分钟内从基础移动进化到复杂几何轨迹。
  • 适合资源受限系统、自主机器人学习与微型智能体开发场景。

本研究探索小型、资源受限系统(如厘米级毫微机器人)如何在其生命周期中自主探索、学习并适应新能力。通过强化学习算法,结合提出的tinyDSM框架,该框架融合内在动机与基于适应度的评估机制,引导智能体进行技能获取与自适应。我们致力于最小化硬编码技能,鼓励开放式的新型技能发展。方法核心在于仅编码极少量先验通用知识,作为系统学习特定依赖关系的基础起点,从而支持广泛的应用领域。该方法基于三要素:(a) 带有内在动机的发展机制;(b) 认知架构(知识、推理、学习);(c) 极低资源消耗。采用分层知识图谱与运动学推理器建模和评估简单及高级运动技能。实验中,使用体积为36 cm³、搭载RP2040微控制器(9 kB内存)的资源受限毫微机器人,集成全部功能(除摄像头外),从最基础的运动技能开始,在15分钟内自主完成从线性/角运动到复杂几何路径的演进。此外,通过仿真分析系统比较不同学习算法与内在动机参数,验证方法有效性。

原文摘要 · Abstract (English)

In this study, we investigate developmental mechanisms that enable small, resource-constrained systems such as cm-sized millirobots to autonomously explore, learn, and adapt their capabilities throughout their lifespan. Reinforcement learning algorithms guide the agent's skill acquisition and adaptation through the interplay of our proposed tinyDSM, which integrates intrinsic motivation and fitness-based assessment. We strive for minimal, hard-wired skills while encouraging the open-ended development of new skills. A key emphasis in our approach is to encode minimal a-priori general knowledge, which serves as a foundational starting point for the system as it further learns system-specific dependencies from the initial knowledge provided. Thus, by design, our approach attempts to cover very generic application domains. The methodology is based on (a) developmental mechanism with intrinsic motivation, and (b) a cognitive architecture (knowledge, reasoning, learning), while (c) utilizing minimal resources. It uses a hierarchical knowledge graph and kinematic reasoners to model and evaluate simple and advanced motion related skills. In our experiments, we use a resource-constrained millirobot with a volume of 36 cm^3 with a Raspberry Pi Pico 32-bit microcontroller (RP2040) that integrates all described features and capabilities except the camera system in 9 kB. Starting with learning the most elementary motor skills the millirobot autonomously progresses from simple linear and angular movements to complex geometric patterns within 15 minutes. To complement the physical experiments, we perform a simulation-based analysis that enables systematic comparisons across learning algorithms and intrinsic motivation parameters.

毫米机器人强化学习技能演化低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。