arXiv:2503.18211cs.CV2025-03CVPR被引 32

通过预测动作相似性提升文本控制动作编辑的精准度

SimMotionEdit: Text-Based Human Motion Editing with Motion Similarity Prediction

  • 联合训练动作编辑与相似性预测,学习语义表征
  • 在MotionFix数据集上实现最高编辑对齐度与保真度
  • 适合需要精确动作控制的动画生成与游戏开发场景

基于文本的3D人体动作编辑是计算机视觉与图形学中的关键挑战任务。尽管无需训练的方法已被探索,但最近发布的MotionFix数据集(包含源文本-动作三元组)为训练提供了新路径,取得了良好效果。然而,现有方法在精确控制方面仍存在不足,常导致动作语义与语言指令不一致。本文提出动作相似性预测这一相关任务,并采用多任务训练范式,让模型在动作编辑与相似性预测中联合学习,以增强语义表征能力。为此,设计了一种基于Diffusion-Transformer的先进架构,分别处理相似性预测与动作编辑。大量实验表明,该方法在编辑对齐度和保真度上均达到当前最优水平。

原文摘要 · Abstract (English)

Text-based 3D human motion editing is a critical yet challenging task in computer vision and graphics. While training-free approaches have been explored, the recent release of the MotionFix dataset, which includes source-text-motion triplets, has opened new avenues for training, yielding promising results. However, existing methods struggle with precise control, often leading to misalignment between motion semantics and language instructions. In this paper, we introduce a related task, motion similarity prediction, and propose a multi-task training paradigm, where we train the model jointly on motion editing and motion similarity prediction to foster the learning of semantically meaningful representations. To complement this task, we design an advanced Diffusion-Transformer-based architecture that separately handles motion similarity prediction and motion editing. Extensive experiments demonstrate the state-of-the-art performance of our approach in both editing alignment and fidelity.

动作编辑文本控制扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。