arXiv:2410.11404cs.CV2024-10被引 6

首个支持多轮对话的细粒度人体运动时空定位模型

MoChat: Joints-Grouped Spatio-Temporal Grounding LLM for Multi-Turn Motion Comprehension and Description

  • 按人体结构分组关节,用编码器融合时空信息
  • 在多个数据集上超越现有方法,实现精准动作定位
  • 适合需要理解复杂动作细节的研究与应用

尽管深度学习在理解人体运动方面持续进步,现有模型仍难以准确识别动作发生的时间和具体身体部位,且通常仅支持单轮交互。为此,我们提出MoChat,一个能够进行人体运动时空定位并理解多轮对话上下文的多模态大语言模型。通过基于人体解剖结构对骨架帧的空间信息进行分组,并引入关节分组骨架编码器,将输出与大语言模型嵌入结合,分别生成具有空间感知和时间感知的嵌入表示。同时,我们构建了基于文本标注提取时间戳的流水线,并生成多轮对话以实现空间定位。最后,通过联合训练多种任务指令,提升模型能力。实验表明,MoChat在多个运动理解任务中达到领先性能,成为首个实现细粒度时空定位的人体运动理解模型。

原文摘要 · Abstract (English)

Despite continuous advancements in deep learning for understanding human motion, existing models often struggle to accurately identify action timing and specific body parts, typically supporting only single-round interaction. Such limitations in capturing fine-grained motion details reduce their effectiveness in motion understanding tasks. In this paper, we propose MoChat, a multimodal large language model capable of spatio-temporal grounding of human motion and understanding multi-turn dialogue context. To achieve these capabilities, we group the spatial information of each skeleton frame based on human anatomical structure and then apply them with Joints-Grouped Skeleton Encoder, whose outputs are combined with LLM embeddings to create spatio-aware and temporal-aware embeddings separately. Additionally, we develop a pipeline for extracting timestamps from skeleton sequences based on textual annotations, and construct multi-turn dialogues for spatially grounding. Finally, various task instructions are generated for jointly training. Experimental results demonstrate that MoChat achieves state-of-the-art performance across multiple metrics in motion understanding tasks, making it as the first model capable of fine-grained spatio-temporal grounding of human motion.

运动理解多轮对话时空定位大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。