arXiv:2601.21904cs.CV2026-01被引 3

通过分层对齐局部动作与文本,提升动作-语言检索精度

Beyond Global Alignment: Fine-Grained Motion-Language Retrieval via Pyramidal Shapley-Taylor Learning

  • 分层建模动作片段与身体关节的细粒度对齐
  • 在多个数据集上显著超越现有方法
  • 适合需要精准动作语义理解的研究者

作为以人为本的跨模态智能基础任务,动作-语言检索旨在弥合自然语言与人体动作之间的语义鸿沟,实现直观的动作分析。然而,现有方法主要关注整个动作序列与全局文本表示的对齐,忽视了局部动作片段、身体关节与文本词元之间的细粒度交互,导致检索性能受限。为此,我们受人类动作感知(从关节动态到片段协同,再到整体理解)的分层过程启发,提出一种新型分层谢尔普利-泰勒(Pyramidal Shapley-Taylor, PST)学习框架,用于细粒度动作-语言检索。该框架将人体动作分解为时间片段和空间身体关节,通过分层的关节级与片段级对齐,逐步学习跨模态对应关系,有效捕捉局部语义细节与层级结构关联。在多个公开基准数据集上的大量实验表明,本方法显著优于当前最优方法,在动作片段、身体关节与其对应文本词元之间实现了精准对齐。代码将在论文被接受后发布。

原文摘要 · Abstract (English)

As a foundational task in human-centric cross-modal intelligence, motion-language retrieval aims to bridge the semantic gap between natural language and human motion, enabling intuitive motion analysis, yet existing approaches predominantly focus on aligning entire motion sequences with global textual representations. This global-centric paradigm overlooks fine-grained interactions between local motion segments and individual body joints and text tokens, inevitably leading to suboptimal retrieval performance. To address this limitation, we draw inspiration from the pyramidal process of human motion perception (from joint dynamics to segment coherence, and finally to holistic comprehension) and propose a novel Pyramidal Shapley-Taylor (PST) learning framework for fine-grained motion-language retrieval. Specifically, the framework decomposes human motion into temporal segments and spatial body joints, and learns cross-modal correspondences through progressive joint-wise and segment-wise alignment in a pyramidal fashion, effectively capturing both local semantic details and hierarchical structural relationships. Extensive experiments on multiple public benchmark datasets demonstrate that our approach significantly outperforms state-of-the-art methods, achieving precise alignment between motion segments and body joints and their corresponding text tokens. The code of this work will be released upon acceptance.

动作理解跨模态检索细粒度对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。