arXiv:2607.22702cs.CVcs.LG2026-07

MIME首次专为双人互动动作设计多模态编码器,提升语言与动作的对齐精度。

MIME: Multimodal Interactive Motion Encoder

论文配图:MIME: Multimodal Interactive Motion Encoder
图 1 · 摘自论文原文
  • 基于流式共注意力与显式交互特征捕捉个体与共享结构
  • 在2000样本库上文本到动作检索准确率提升12.8%相对性能
  • 可作为冻结先验迁移至其他生成模型,提升语义对齐效果

文本-动作表征学习快速发展,尤其在动画、AR/VR和具身AI中对多人交互建模需求增长。此类场景需同时对齐语言与个体动作及人物间关系。我们提出多模态交互动作编码器(MIME),据我们所知是首个专为双人互动动作设计的多模态编码器。MIME通过基于流的共注意力机制与显式交互特征,捕捉个体与共享结构,并采用课程式对比训练。在Inter-X文本-动作检索任务中,MIME在不同候选集规模下均优于早期融合与晚期融合基线,在2,000样本库上实现文本到动作检索的R@1提升12.8%相对性能。进一步评估显示,将MIME作为冻结辅助先验引入TIMotion与InterMask,在未见的InterHuman数据集上显著提升语义对齐指标,同时保持相近的FID值。结果表明,具备交互感知能力的多模态编码能有效提升多人动作检索性能,并支持跨数据集迁移以服务下游动作生成任务。

原文摘要 · Abstract (English)

Text-motion representation learning has advanced rapidly, with growing interest in multi person interactions for animation, AR/VR, and embodied AI. These settings require representations that align language with both individual actor dynamics and the relationships between actors. We introduce the Multimodal Interactive Motion Encoder (MIME), which, to our knowledge, represents the first dedicated multimodal encoder designed specifically for two person interactive motion. MIME captures individual and shared structure using stream based co-attention with explicit interaction features and curriculum based contrastive training. On Inter-X text-motion retrieval, MIME consistently outperforms early and late fusion baselines across gallery sizes, achieving a 12.8% relative improvement in text-to-motion R@1 at a 2,000-sample gallery. We further evaluate MIME as a frozen auxiliary prior within TIMotion and InterMask on the unseen InterHuman dataset. MIME improves semantic alignment metrics while maintaining comparable FID in TIMotion. These results show that interaction aware multimodal encoding improves multi person motion retrieval and transfers across datasets to support downstream motion generation.

动作生成多模态交互建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。