arXiv:2502.18180cs.AIcs.MA2025-02被引 16

ChatMotion用多智能体实现人机动态分析,更懂用户需求。

ChatMotion: A Multimodal Multi-Agent for Human Motion Analysis

  • 构建多智能体框架,动态理解意图并拆解任务
  • 集成运动核心模块,多角度分析人机动态
  • 适合需要交互式动作分析的研究与应用

多模态大语言模型在人运动理解方面取得进展,但其仅支持指令输入,缺乏交互性与多视角适应能力。为此,我们提出ChatMotion,一个用于人运动分析的多模态多智能体框架。该框架可动态解析用户意图,将复杂任务分解为元任务,并激活专用功能模块进行运动理解。系统集成多个专用模块(如MotionCore),从不同角度分析人运动。大量实验表明,ChatMotion在人运动理解中具备高精度、强适应性和良好用户参与度。

原文摘要 · Abstract (English)

Advancements in Multimodal Large Language Models (MLLMs) have improved human motion understanding. However, these models remain constrained by their "instruct-only" nature, lacking interactivity and adaptability for diverse analytical perspectives. To address these challenges, we introduce ChatMotion, a multimodal multi-agent framework for human motion analysis. ChatMotion dynamically interprets user intent, decomposes complex tasks into meta-tasks, and activates specialized function modules for motion comprehension. It integrates multiple specialized modules, such as the MotionCore, to analyze human motion from various perspectives. Extensive experiments demonstrate ChatMotion's precision, adaptability, and user engagement for human motion understanding.

人运动分析多智能体多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。