arXiv:2507.23188cs.CV2025-07中稿 · IEEE TMM 2025被引 3

四模态对齐实现更精准的动捕检索,提升沉浸感与交互便利性。

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space

  • 通过序列级对比学习构建细粒度联合嵌入空间,对齐文本、音频、视频与动作
  • 在HumanML3D上实现文本到动作检索R@10提升10.16%,视频到动作检索R@1提升25.43%
  • 首次引入音频模态,适合需要高精度动作检索的应用场景

动作检索对于动作获取至关重要,相比动作生成具有更高的精度、真实感、可控性和可编辑性。现有方法利用对比学习构建统一嵌入空间,实现从文本或视觉模态进行动作检索。然而,这些方法缺乏直观友好的交互方式,且常忽略多数模态的时序特性,影响检索性能。为此,我们提出一个框架,将文本、音频、视频和动作四种模态对齐至细粒度联合嵌入空间,首次在动作检索中引入音频以增强用户沉浸感与操作便捷性。该细粒度空间通过序列级对比学习实现,能捕捉跨模态关键细节以促进更好对齐。为评估框架,我们在现有文本-动作数据集基础上合成多样化音频,构建两个多模态动作检索数据集。实验结果表明,在多个子任务上均优于现有最优方法:在HumanML3D数据集上,文本到动作检索的R@10提升10.16%,视频到动作检索的R@1提升25.43%。此外,四模态框架显著优于三模态版本,凸显多模态动作检索在动作获取中的潜力。

原文摘要 · Abstract (English)

Motion retrieval is crucial for motion acquisition, offering superior precision, realism, controllability, and editability compared to motion generation. Existing approaches leverage contrastive learning to construct a unified embedding space for motion retrieval from text or visual modality. However, these methods lack a more intuitive and user-friendly interaction mode and often overlook the sequential representation of most modalities for improved retrieval performance. To address these limitations, we propose a framework that aligns four modalities -- text, audio, video, and motion -- within a fine-grained joint embedding space, incorporating audio for the first time in motion retrieval to enhance user immersion and convenience. This fine-grained space is achieved through a sequence-level contrastive learning approach, which captures critical details across modalities for better alignment. To evaluate our framework, we augment existing text-motion datasets with synthetic but diverse audio recordings, creating two multi-modal motion retrieval datasets. Experimental results demonstrate superior performance over state-of-the-art methods across multiple sub-tasks, including an 10.16% improvement in R@10 for text-to-motion retrieval and a 25.43% improvement in R@1 for video-to-motion retrieval on the HumanML3D dataset. Furthermore, our results show that our 4-modal framework significantly outperforms its 3-modal counterpart, underscoring the potential of multi-modal motion retrieval for advancing motion acquisition.

动作检索多模态嵌入空间音频对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。