用半在线偏好优化提升文本生成动作的自然度与一致性。
SoPo: Text-to-Motion Generation Using Semi-Online Preference Optimization
- 结合在线与离线数据,用半在线方式优化动作生成偏好。
- 在MLD和MDM模型上,运动质量指标分别提升3.25%和2.91%。
- 适合追求高质量动作生成的创作者与动画研发人员。
文本到动作生成对创意产业至关重要,但常难以生成一致且逼真的动作。本文聚焦于微调文本到动作模型,使其稳定偏好高质量、受人类喜爱的动作,这是关键但尚未深入研究的问题。我们理论上分析了在线与离线偏好优化(DPO)的局限性:离线DPO易过拟合,在线DPO存在采样偏差。基于此,提出半在线偏好优化(SoPo),利用由在线分布中非优选动作和离线数据集中优选动作组成的数据对进行训练。该方法融合在线与离线DPO优势,相互弥补缺陷。大量实验表明,SoPo优于其他偏好对齐方法:在MLD模型上MM-Dist为3.25%(如MoDiPO为0.76%),在MDM模型上为2.91%(如MoDiPO为0.66%)。此外,经SoPo微调的MLD模型在R-precision和MM-Dist上超越当前最优模型。可视化结果也验证了其在偏好对齐上的有效性。
原文摘要 · Abstract (English)
Text-to-motion generation is essential for advancing the creative industry but often presents challenges in producing consistent, realistic motions. To address this, we focus on fine-tuning text-to-motion models to consistently favor high-quality, human-preferred motions, a critical yet largely unexplored problem. In this work, we theoretically investigate the DPO under both online and offline settings, and reveal their respective limitation: overfitting in offline DPO, and biased sampling in online DPO. Building on our theoretical insights, we introduce Semi-online Preference Optimization (SoPo), a DPO-based method for training text-to-motion models using "semi-online" data pair, consisting of unpreferred motion from online distribution and preferred motion in offline datasets. This method leverages both online and offline DPO, allowing each to compensate for the other's limitations. Extensive experiments demonstrate that SoPo outperforms other preference alignment methods, with an MM-Dist of 3.25% (vs e.g. 0.76% of MoDiPO) on the MLD model, 2.91% (vs e.g. 0.66% of MoDiPO) on MDM model, respectively. Additionally, the MLD model fine-tuned by our SoPo surpasses the SoTA model in terms of R-precision and MM Dist. Visualization results also show the efficacy of our SoPo in preference alignment. Project page: https://xiaofeng-tan.github.io/projects/SoPo/ .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。