arXiv:2508.03343cs.CV2025-08被引 7

用小波分析人体运动多频特征,提升文本与动作匹配精度。

WaMo: Wavelet-Enhanced Multi-Frequency Trajectory Analysis for Fine-Grained Text-Motion Retrieval

  • 通过小波分解捕捉关节轨迹的多尺度运动细节。
  • 在HumanML3D和KIT-ML上分别提升17.0%和18.2%检索性能。
  • 适合需要精细动作语义对齐的研究者或开发者。

文本-动作检索(TMR)旨在根据文本描述检索语义相关的3D动作序列。然而,由于人体结构复杂且具有时空动态特性,实现动作与文本的精准匹配仍具挑战。现有方法常忽略这些细节,仅采用通用编码方式,难以区分不同身体部位及其动态变化,限制了语义对齐精度。为此,本文提出WaMo,一种基于小波的多频特征提取框架,可从个体关节轨迹中多分辨率地捕获特定关节与时间变化的运动细节,提取具有判别性的运动特征以实现细粒度对齐。WaMo包含三个核心组件:(1) 运动轨迹小波分解,将运动信号分解为保留局部运动细节与全局语义的频率分量;(2) 小波重构模块,利用可学习的逆小波变换从提取特征中重建原始关节轨迹,确保关键时空信息不丢失;(3) 无序动作序列预测,对打乱的动作序列进行重排序,强化对内在时序一致性的学习,提升动作-文本对齐效果。大量实验表明,WaMo在HumanML3D和KIT-ML数据集上分别取得17.0%和18.2%的相对提升,优于当前最优方法。代码已开源:https://github.com/3DAgentWorld/WaMo/

原文摘要 · Abstract (English)

Text-Motion Retrieval (TMR) aims to retrieve 3D motion sequences semantically relevant to text descriptions. However, matching 3D motions with text remains highly challenging, primarily due to the intricate structure of the human body and its spatiotemporal dynamics. Existing approaches often overlook these complexities, relying on general encoding methods that fail to distinguish different body parts and their dynamics, limiting precise semantic alignment. To address this, we propose WaMo, a novel wavelet-based multi-frequency feature extraction framework. It fully captures joint-specific and time-varying motion details at multiple resolutions for individual joint trajectories, extracting discriminative motion features to achieve fine-grained alignment with texts. WaMo has three key components: (1) Trajectory Wavelet Decomposition decomposes motion signals into frequency components that preserve both local kinematic details and global motion semantics. (2) Trajectory Wavelet Reconstruction uses learnable inverse wavelet transforms to reconstruct original joint trajectories from extracted features, ensuring the preservation of essential spatiotemporal information. (3) Disordered Motion Sequence Prediction reorders shuffled motion sequences to improve learning of inherent temporal coherence, enhancing motion-text alignment. Extensive experiments demonstrate WaMo's superiority, achieving 17.0\% and 18.2\% relative improvements in $Rsum$ on HumanML3D and KIT-ML datasets, respectively, outperforming existing state-of-the-art (SOTA) methods. Code is available at https://github.com/3DAgentWorld/WaMo/.

动作生成小波分析文本对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。