arXiv:2411.18654cs.CV2024-11CVPR被引 9

用GPT-4Vision做奖励,让文本生成动作更贴合事件层次描述。

AToM: Aligning Text-to-Motion Model at Event-Level with GPT-4Vision Reward

  • 用GPT-4Vision为动作生成提供细粒度标注与评分。
  • 在运动完整性、时间关系、频率上提升对齐效果。
  • 适合需要精准动作控制的动画与虚拟人应用。

近期,文本到动作模型为高效灵活地生成真实人体动作开辟了新可能。然而,将动作生成与事件级文本描述对齐面临独特挑战,因文本提示与期望动作结果间存在复杂关系。为此,我们提出AToM框架,通过利用GPT-4Vision的奖励信号增强生成动作与文本提示的对齐性。AToM包含三个阶段:首先构建MotionPrefer数据集,将三类事件级文本提示与生成动作配对,覆盖动作完整性、时间关系与频率;其次设计基于GPT-4Vision的详细动作标注范式,包括视觉数据格式化、任务特定指令及各子任务评分规则;最后使用该范式引导的强化学习微调现有文本到动作模型。实验表明,AToM显著提升了文本到动作生成在事件级对齐上的质量。

原文摘要 · Abstract (English)

Recently, text-to-motion models have opened new possibilities for creating realistic human motion with greater efficiency and flexibility. However, aligning motion generation with event-level textual descriptions presents unique challenges due to the complex relationship between textual prompts and desired motion outcomes. To address this, we introduce AToM, a framework that enhances the alignment between generated motion and text prompts by leveraging reward from GPT-4Vision. AToM comprises three main stages: Firstly, we construct a dataset MotionPrefer that pairs three types of event-level textual prompts with generated motions, which cover the integrity, temporal relationship and frequency of motion. Secondly, we design a paradigm that utilizes GPT-4Vision for detailed motion annotation, including visual data formatting, task-specific instructions and scoring rules for each sub-task. Finally, we fine-tune an existing text-to-motion model using reinforcement learning guided by this paradigm. Experimental results demonstrate that AToM significantly improves the event-level alignment quality of text-to-motion generation.

动作生成文本到动作GPT-4Vision对齐优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。