零样本识别动作细粒度子阶段,无需标注直接理解复杂人体行为。
Zero-Shot Open-Vocabulary Human Motion Grounding with Test-Time Training
- 用大模型拆解语言指令为有序子动作单元,实现语义对齐。
- 通过软掩码优化定位关键帧,提升分割连续性与段间区分度。
- 无需微调即可在真实场景中应用,适合无标注动作理解任务。
理解复杂人体活动需要将运动分解为细粒度、语义对齐的子动作。这一运动定位过程对行为分析、具身智能和虚拟现实至关重要。然而,现有方法多依赖预定义动作类别的密集标注,难以适应开放词汇、真实世界场景。本文提出ZOMG,一种零样本、开放词汇框架,可在不依赖任何标注或微调的情况下,将运动序列分割为语义有意义的子动作。技术上,ZOMG结合(1)语言语义分割,利用大语言模型将指令分解为有序子动作单元;(2)软掩码优化,学习实例特定的时间掩码,聚焦于子动作的关键帧,同时保持段内连续性和段间分离性,且不修改预训练编码器。在三个运动-语言数据集上的实验表明,该方法在运动定位性能上达到当前最佳,相较于之前方法在HumanML3D基准上提升+8.7% mAP。同时,在下游检索任务中也取得显著改进,确立了无标注运动理解的新范式。
原文摘要 · Abstract (English)
Understanding complex human activities demands the ability to decompose motion into fine-grained, semantic-aligned sub-actions. This motion grounding process is crucial for behavior analysis, embodied AI and virtual reality. Yet, most existing methods rely on dense supervision with predefined action classes, which are infeasible in open-vocabulary, real-world settings. In this paper, we propose ZOMG, a zero-shot, open-vocabulary framework that segments motion sequences into semantically meaningful sub-actions without requiring any annotations or fine-tuning. Technically, ZOMG integrates (1) language semantic partition, which leverages large language models to decompose instructions into ordered sub-action units, and (2) soft masking optimization, which learns instance-specific temporal masks to focus on frames critical to sub-actions, while maintaining intra-segment continuity and enforcing inter-segment separation, all without altering the pretrained encoder. Experiments on three motion-language datasets demonstrate state-of-the-art effectiveness and efficiency of motion grounding performance, outperforming prior methods by +8.7\% mAP on HumanML3D benchmark. Meanwhile, significant improvements also exist in downstream retrieval, establishing a new paradigm for annotation-free motion understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。