让机器人实时生成带情绪的自然动作,用数值参数控制情感强度。
PAMoR: Parameterized Affective Motion Generation in Real Time for Humanoid Robots

- 通过姿态扩展和运动能量计算情感坐标,无需人工标注。
- 在29自由度机器人上实现动作与情感同步生成,支持实时编辑。
- 情感识别准确率达0.38,接近真人表演水平。
人们在社交场景中不仅关注人形机器人的动作本身,也关注其传达的情感。现有方法生成情感动作依赖参考视频或情感词,难以量化。本文提出PAMoR,将情感转化为可测量的控制参数:基于机器人本体运动学的价-唤醒(V-A)坐标,通过姿态展开和运动能量闭式求解获得,直接作为生成条件,无需人工标注。在共享潜在空间中训练的动作先验与两个情感先验,在去噪每一步组合使用:动作先验决定执行内容,情感先验调节表现方式。整个身体动作在29自由度的Unitree G1机器人上实时自回归生成,动作与情感均可编辑。生成动作能覆盖全范围的指令V-A值,且文本到动作的匹配度不逊于纯文本基线。感知实验显示,评分者在0.38的试验中正确识别了指令情感,高于基线并接近真人表演的0.44。
原文摘要 · Abstract (English)
People read a humanoid robot's motion in social settings not only for the action performed but for the affect conveyed. Motion carrying that affect has so far been generated for human avatars, where style is taken from a reference clip or an emotion word, neither of which can be quantitatively parameterized. We present PAMoR, which turns affect into a measured control parameter: a valence-arousal (V-A) coordinate computed natively on robot kinematics. It is obtained in closed form from postural expansion and movement energy, and these measurements serve directly as generation conditions, with no human annotation. An action prior and two affect priors, trained in a shared latent space, are composed at each denoising step: the action prior fixes what is performed, the affect priors modulate how. Whole-body motion rolls out autoregressively on a 29-DoF Unitree G1 in real time, with action and affect both editable. Generated motion tracks the commanded V-A over its full range while text-to-motion fidelity still matches text-only baselines. In a perceptual study, raters identify the commanded emotion on 0.38 of trials, above both baselines and approaching the 0.44 reported for acted human bodies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。