让虚拟人物互动更真实,能保留个人特征并精准匹配文本描述。
InterMoE: Individual-Specific 3D Human Interaction Generation via Dynamic Temporal-Selective MoE
- 用动态路由机制结合语义与动作上下文分配任务给不同专家。
- 在InterHuman和InterX数据集上FID分别降低9%和22%,效果领先。
- 适合需要个性化角色互动的VR、游戏和机器人应用。
生成高质量的人类交互对虚拟现实和机器人等应用具有重要意义。然而,现有方法常无法保持个体独特特征或完全遵循文本描述。为此,我们提出InterMoE,一种基于动态时序选择性专家混合(Dynamic Temporal-Selective Mixture of Experts)的新框架。其核心是路由机制,协同利用高层文本语义与底层运动上下文,将时序动作特征分发给专业专家。该机制使专家可动态调整选择能力,聚焦关键时序特征,从而在保证高语义保真度的同时,有效保留个体特征身份。大量实验表明,InterMoE在个体特异性高保真3D人体交互生成任务中达到当前最优性能,在InterHuman数据集上FID降低9%,在InterX数据集上降低22%。
原文摘要 · Abstract (English)
Generating high-quality human interactions holds significant value for applications like virtual reality and robotics. However, existing methods often fail to preserve unique individual characteristics or fully adhere to textual descriptions. To address these challenges, we introduce InterMoE, a novel framework built on a Dynamic Temporal-Selective Mixture of Experts. The core of InterMoE is a routing mechanism that synergistically uses both high-level text semantics and low-level motion context to dispatch temporal motion features to specialized experts. This allows experts to dynamically determine the selection capacity and focus on critical temporal features, thereby preserving specific individual characteristic identities while ensuring high semantic fidelity. Extensive experiments show that InterMoE achieves state-of-the-art performance in individual-specific high-fidelity 3D human interaction generation, reducing FID scores by 9% on the InterHuman dataset and 22% on InterX.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。