arXiv:2410.06513cs.CV2024-10被引 14

用多奖励强化学习让文本生成动作更符合人类偏好。

MotionRL: Align Text-to-Motion Generation to Human Preferences with Multi-Reward Reinforcement Learning

  • 通过人类偏好先验知识指导动作生成的强化学习优化
  • 在文本匹配度、动作质量和人类偏好上均显著提升
  • 支持多目标权衡,适合需要人性化动作生成的应用

我们提出MotionRL,首个将多奖励强化学习用于文本到动作生成并对其对齐人类偏好的方法。以往工作仅关注数据集上的数值指标,忽视了人类反馈的多样性和主观性。MotionRL利用强化学习,基于人类感知模型的先验知识微调动作生成器,使生成动作更契合人类偏好。同时引入新型多目标优化策略,逼近文本一致性、动作质量与人类偏好之间的帕累托最优。大量实验与用户研究显示,MotionRL不仅可灵活控制不同目标下的生成结果,且在各项指标上显著优于现有算法。

原文摘要 · Abstract (English)

We introduce MotionRL, the first approach to utilize Multi-Reward Reinforcement Learning (RL) for optimizing text-to-motion generation tasks and aligning them with human preferences. Previous works focused on improving numerical performance metrics on the given datasets, often neglecting the variability and subjectivity of human feedback. In contrast, our novel approach uses reinforcement learning to fine-tune the motion generator based on human preferences prior knowledge of the human perception model, allowing it to generate motions that better align human preferences. In addition, MotionRL introduces a novel multi-objective optimization strategy to approximate Pareto optimality between text adherence, motion quality, and human preferences. Extensive experiments and user studies demonstrate that MotionRL not only allows control over the generated results across different objectives but also significantly enhances performance across these metrics compared to other algorithms.

动作生成强化学习人类偏好

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。