arXiv:2512.23077cs.RO2025-12

用视觉语言模型自动学习复杂动作的奖励函数,让机器人更像人一样自然运动。

Embodied Learning of Reward for Musculoskeletal Control with Vision Language Models

论文配图:Embodied Learning of Reward for Musculoskeletal Control with Vision Language Models
图 1 · 摘自论文原文
  • 利用视觉语言模型动态生成奖励信号,替代人工设计。
  • 在高维肌肉骨骼系统上实现稳定行走与姿态控制。
  • 适合研究具身智能、动作生成与人机交互的学者。

高维肌肉骨骼系统的运动控制中,有效奖励函数的设计仍是核心挑战。人类能明确描述运动目标,如“保持直立姿势向前行走”,但实现这些目标的控制策略大多隐含难显。本文提出运动视觉语言表征(MoVLR)框架,利用视觉语言模型(VLMs)将高层目标与运动控制相连接。该方法通过控制优化与VLM反馈的迭代交互,自主探索奖励空间,使控制策略与物理协调行为对齐。本方法将视觉与语言评估转化为结构化指导,实现了高维肌肉骨骼运动与操作任务中奖励函数的发现与优化。结果表明,视觉语言模型可有效将抽象运动描述锚定于生理运动控制的隐含规律。

原文摘要 · Abstract (English)

Discovering effective reward functions remains a fundamental challenge in motor control of high-dimensional musculoskeletal systems. While humans can describe movement goals explicitly such as "walking forward with an upright posture," the underlying control strategies that realize these goals are largely implicit, making it difficult to directly design rewards from high-level goals and natural language descriptions. We introduce Motion from Vision-Language Representation (MoVLR), a framework that leverages vision-language models (VLMs) to bridge the gap between goal specification and movement control. Rather than relying on handcrafted rewards, MoVLR iteratively explores the reward space through iterative interaction between control optimization and VLM feedback, aligning control policies with physically coordinated behaviors. Our approach transforms language and vision-based assessments into structured guidance for embodied learning, enabling the discovery and refinement of reward functions for high-dimensional musculoskeletal locomotion and manipulation. These results suggest that VLMs can effectively ground abstract motion descriptions in the implicit principles governing physiological motor control.

具身智能视觉语言模型运动控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。