用强化学习提升多模态模型生成3D人体姿态的准确性
Pose-RFT: Enhancing MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-Tuning
- 将姿态生成设为离散语言与连续姿态的混合强化学习任务
- 在多个基准上显著超越现有模型,图像到姿态误差降低12.3%
- 适合需要精准姿态生成的动画、VR场景开发者
从图像或文本等多模态输入生成3D人体姿态,要求模型同时捕捉丰富的空间与语义对应关系。尽管针对姿态的多模态大语言模型(MLLM)已展现出潜力,但其通常采用监督学习目标(如SMPL参数回归或标记级预测),难以建模固有的模糊性,也难以实现准确生成所需的特定任务对齐。为此,我们提出Pose-RFT,一种专用于MLLM中3D人体姿态生成的强化微调框架。我们将该任务建模为混合动作强化学习问题,联合优化离散语言预测与连续姿态生成。为此,我们引入HyGRPO算法,通过采样响应的组内奖励归一化,指导离散与连续动作的联合优化。Pose-RFT进一步引入任务特异性奖励函数,分别引导图像到姿态生成中的空间对齐与文本到姿态生成中的语义一致性。在多个姿态生成基准上的大量实验表明,Pose-RFT显著优于现有姿态专用MLLM,验证了混合动作强化微调在3D姿态生成中的有效性。
原文摘要 · Abstract (English)
Generating 3D human poses from multimodal inputs such as images or text requires models to capture both rich spatial and semantic correspondences. While pose-specific multimodal large language models (MLLMs) have shown promise in this task, they are typically trained with supervised objectives such as SMPL parameter regression or token-level prediction, which struggle to model the inherent ambiguity and achieve task-specific alignment required for accurate 3D pose generation. To address these limitations, we propose Pose-RFT, a reinforcement fine-tuning framework tailored for 3D human pose generation in MLLMs. We formulate the task as a hybrid action reinforcement learning problem that jointly optimizes discrete language prediction and continuous pose generation. To this end, we introduce HyGRPO, a hybrid reinforcement learning algorithm that performs group-wise reward normalization over sampled responses to guide joint optimization of discrete and continuous actions. Pose-RFT further incorporates task-specific reward functions to guide optimization towards spatial alignment in image-to-pose generation and semantic consistency in text-to-pose generation. Extensive experiments on multiple pose generation benchmarks demonstrate that Pose-RFT significantly improves performance over existing pose-specific MLLMs, validating the effectiveness of hybrid action reinforcement fine-tuning for 3D pose generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。