arXiv:2507.19850cs.CV2025-07ICCV被引 9

构建细粒度人体运动数据集,支持文本驱动的精准动作生成与编辑

FineMotion: A Dataset and Benchmark with both Spatial and Temporal Annotation for Fine-grained Motion Generation and Editing

  • 构建含44.2万段动作片段的细粒度人体运动数据集
  • 在MDM模型上实现15.3%的Top-3准确率提升
  • 支持零样本下空间与时间维度的文本控制动作编辑

从文本描述生成逼真人体动作已取得显著进展,但现有方法常忽略特定身体部位及其运动时序。本文通过丰富文本描述细节来解决该问题,提出FineMotion数据集,包含超过44.2万段人体动作片段及其对应的身体部位运动详细描述,另含约9.5万段完整动作序列的身体部位运动描述段落。实验表明,该数据集对文本驱动的细粒度人体动作生成任务具有重要意义,尤其使MDM模型在Top-3准确率上提升了15.3%。此外,我们还实现了无需训练的细粒度动作编辑零样本流程,支持通过文本在空间和时间维度进行精确编辑。数据集与代码已公开于CVI-SZU/FineMotion。

原文摘要 · Abstract (English)

Generating realistic human motions from textual descriptions has undergone significant advancements. However, existing methods often overlook specific body part movements and their timing. In this paper, we address this issue by enriching the textual description with more details. Specifically, we propose the FineMotion dataset, which contains over 442,000 human motion snippets - short segments of human motion sequences - and their corresponding detailed descriptions of human body part movements. Additionally, the dataset includes about 95k detailed paragraphs describing the movements of human body parts of entire motion sequences. Experimental results demonstrate the significance of our dataset on the text-driven finegrained human motion generation task, especially with a remarkable +15.3% improvement in Top-3 accuracy for the MDM model. Notably, we further support a zero-shot pipeline of fine-grained motion editing, which focuses on detailed editing in both spatial and temporal dimensions via text. Dataset and code available at: CVI-SZU/FineMotion

动作生成细粒度数据集文本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。