arXiv:2411.16805cs.AIcs.CV2024-11CVPR被引 33

让模型直接理解原始动作数据,提升对复杂行为的分析能力

Human Motion Instruction Tuning

  • 保留动作原始形式进行指令微调,避免信息丢失
  • 在体育与专业活动场景中显著提升行为理解与预测性能
  • 适合需要精准动作分析的应用,如运动分析与行为预测

本文提出一种名为LLaMo(Large Language and Human Motion Assistant)的多模态框架,用于人类动作的指令微调。与传统方法将视频或动作序列转为语言标记不同,LLaMo在指令微调过程中保持动作数据的原始形式,从而保留动作特有的细节,提升模型对复杂人类行为的理解能力。通过同时处理视频、动作数据和文本输入,LLaMo实现灵活且以人为核心的分析。在高复杂度领域(包括人类行为与专业活动)的实验评估表明,该模型能有效捕捉领域知识,在动作密集型场景中增强理解与预测能力。我们希望LLaMo能为未来的多模态人工智能系统提供基础,广泛应用于体育分析与行为预测等领域。代码与模型已在项目官网公开:https://github.com/ILGLJ/LLaMo。

原文摘要 · Abstract (English)

This paper presents LLaMo (Large Language and Human Motion Assistant), a multimodal framework for human motion instruction tuning. In contrast to conventional instruction-tuning approaches that convert non-linguistic inputs, such as video or motion sequences, into language tokens, LLaMo retains motion in its native form for instruction tuning. This method preserves motion-specific details that are often diminished in tokenization, thereby improving the model's ability to interpret complex human behaviors. By processing both video and motion data alongside textual inputs, LLaMo enables a flexible, human-centric analysis. Experimental evaluations across high-complexity domains, including human behaviors and professional activities, indicate that LLaMo effectively captures domain-specific knowledge, enhancing comprehension and prediction in motion-intensive scenarios. We hope LLaMo offers a foundation for future multimodal AI systems with broad applications, from sports analytics to behavioral prediction. Our code and models are available on the project website: https://github.com/ILGLJ/LLaMo.

动作理解多模态指令微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。