让两个人的3D动作按文字指令自然互动,效果领先。
InterEdit: Navigating Text-Guided 3D Dyadic Human Motion Editing
- 用可学习标记捕捉交互语义,通过频域分析建模动作周期性。
- 在多人体动作编辑任务中实现最高文本一致性与编辑保真度。
- 适合研究多人动作生成、人机交互或虚拟角色动画的开发者。
文本引导的3D动作编辑在单人场景中已取得进展,但扩展到多人场景受限于配对数据稀缺及人际交互复杂性。本文提出多人3D动作编辑任务,即根据源动作和文本指令生成目标动作。为此,我们构建了InterEdit3D数据集,包含人工标注的双人动作变化信息,并设立文本引导多人动作编辑(TMME)基准。提出InterEdit模型,一种同步的无分类器条件扩散模型,引入语义感知计划标记对齐(含可学习标记)以捕捉高层交互线索,并设计交互感知频域标记对齐策略,结合DCT与能量池化建模周期性运动动态。实验表明,该模型显著提升文本到动作的一致性和编辑保真度,在TMME任务中达到当前最优性能。数据集与代码将开源于https://github.com/YNG916/InterEdit。
原文摘要 · Abstract (English)
Text-guided 3D motion editing has seen success in single-person scenarios, but its extension to multi-person settings is less explored due to limited paired data and the complexity of inter-person interactions. We introduce the task of multi-person 3D motion editing, where a target motion is generated from a source and a text instruction. To support this, we propose InterEdit3D, a new dataset with manual two-person motion change annotations, and a Text-guided Multi-human Motion Editing (TMME) benchmark. We present InterEdit, a synchronized classifier-free conditional diffusion model for TMME. It introduces Semantic-Aware Plan Token Alignment with learnable tokens to capture high-level interaction cues and an Interaction-Aware Frequency Token Alignment strategy using DCT and energy pooling to model periodic motion dynamics. Experiments show that InterEdit improves text-to-motion consistency and edit fidelity, achieving state-of-the-art TMME performance. The dataset and code will be released at https://github.com/YNG916/InterEdit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。