让3D人体动作编辑更精准可控,支持多种语义级修改。
UniMoFlow: Grounding Instruction-Driven 3D Human Motion Editing in Generation

- 通过闭环合成验证构建大规模动作编辑数据集
- 提出统一流匹配模型实现生成与编辑共享知识
- 支持细粒度动作修改,适合动画、游戏开发应用
指令驱动的3D人体动作编辑需要精确的时空定位、丰富的语义支撑以及未修改内容的严格保留。现有方法或依赖无需训练的生成模型适配,或仅依靠三元组监督;但适配常导致控制效果不佳,而人工标注的三元组数据集规模和语义多样性严重受限。为此,我们直接在文本到动作生成的过程中实现跨数据、架构与推理层面的动作编辑对齐。在数据层面,构建闭环合成与验证流程,生成涵盖身体部位、幅度、时间、动作和风格等多维度编辑的Omni-MoEdit大规模数据集。在架构层面,提出UniMoFlow——一种统一的潜在流匹配模型,使生成与编辑共享广泛的语义与运动学知识。在推理层面,引入SAFE(Source-Anchored Flow Editing)实现可控、源锚定的精细化调整。此外,采用语义感知评估指标,以衡量那些因合理语义变化而偏离单一真实参考的编辑有效性。大量实验表明,该方法在目标文本对齐性、编辑有效性与循环一致性方面均有提升,同时保持了良好的源动作保真度与文本到动作生成质量。
原文摘要 · Abstract (English)
Instruction-driven editing of 3D human motion requires precise spatiotemporal localization, rich semantic grounding, and strict preservation of unmodified content. Existing methods either resort to training-free adaptation of generative models or rely solely on triplet supervision; however, adaptation often yields suboptimal control, and manually curated triplet datasets remain severely limited in scale and semantic diversity. To overcome this bottleneck, we ground motion editing directly within text-to-motion generation across data, architecture, and inference. At the data level, we develop a closed-loop synthesis-and-verification pipeline that produces Omni-MoEdit, a large-scale dataset spanning body-part, amplitude, temporal, action, and style edits. At the architectural level, we introduce UniMoFlow, a unified latent flow-matching model that shares broad semantic and kinematic knowledge between generation and editing. At the inference level, SAFE (Source-Anchored Flow Editing) complements UniMoFlow with controllable, source-anchored refinement. Furthermore, we augment standard evaluations with semantics-aware metrics to account for valid edits that inherently deviate from a single ground-truth reference. Extensive experiments demonstrate improved target-text alignment, edit effectiveness, and cycle consistency, while maintaining competitive source fidelity and text-to-motion generation quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。