用变压器模型让机器人同时掌握多种动作技能。
Arnold: A multi-task, multi-embodiment muscle transformer policy
- 基于变压器架构,统一处理不同任务与身体形态的控制。
- 在14项任务中表现媲美甚至超越单一技能专家模型。
- 适合研究多任务控制与生物运动机制的学者。
控制高维非线性的人体肌肉骨骼模型是基础科学挑战。近年来机器学习突破使得虚拟代理能在多个自由度的肌肉骨骼系统中掌握单个技能,如抓取、物体操作和行走。然而这些代理仅为“专家”,仅擅长单一技能。本文提出Arnold,一种基于变压器的肌肉骨骼控制策略,可同时掌握多项任务与不同身体形态。Arnold结合行为克隆与强化学习,解决14项复杂控制任务,涵盖灵巧操作、伸手及行走,性能匹配或超过单任务专家模型。核心创新在于其传感器-运动词汇表,一种对异构感知模态、目标与执行器语义的组合表示。该词汇表通过变压器架构处理任务间变化的观测与动作空间,支持高效多任务、多形态学习,并促进快速适应新任务,同时鼓励通用运动策略如动作与运动学平滑性。因果探测分析显示,低维肌群协同仍高度任务特异,方差分析系统低估了功能控制维度,与生物观察一致,表明此类协同转移性有限。代码与数据见:https://github.com/amathislab/arnold。
原文摘要 · Abstract (English)
Controlling high-dimensional and nonlinear musculoskeletal models of the human body is a foundational scientific challenge. Recent machine learning breakthroughs have heralded in-silico policies that master individual skills like reaching, object manipulation and locomotion in musculoskeletal systems with many degrees of freedom. However, these agents are merely "specialists", achieving high performance for a single skill. In this work, we develop Arnold, a transformer-based musculoskeletal control policy that masters multiple tasks and embodiments. Arnold combines behavior cloning and reinforcement learning to address 14 challenging control tasks spanning dexterous object manipulation, reaching, and locomotion, matching or exceeding the performance of single-task specialist policies. A key innovation is Arnold's sensorimotor vocabulary, a compositional representation of the semantics of heterogeneous sensory modalities, objectives, and actuators. Arnold leverages this vocabulary via a transformer architecture to deal with the variable observation and action spaces across tasks. This framework supports efficient multi-task, multi-embodiment learning and facilitates rapid adaptation to novel tasks, while encouraging universal motor strategies such as action and kinematic smoothness. Finally, causal probing of the motor output reveals that low-dimensional muscle synergies remain largely task-specific and that variance-based analyses systematically underestimate functional control dimensionality, consistent with biological observations on the limited transferability of such synergies. Code and data are available here: https://github.com/amathislab/arnold
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。