融合动作与视频的多模态模型,提升行为理解准确率
ViMoNet: A Multimodal Vision-Language Framework for Human Behavior Understanding from Motion and Video
- 用动作+视频双模态数据联合训练,捕捉动作细节与上下文语义
- 在多个任务上超越现有方法,尤其在行为解释和描述生成上表现突出
- 适合智能医疗场景,如老人跌倒检测与健康风险预警
本研究探索利用大语言模型(LLMs)通过融合动作与视频数据实现人类行为理解。我们认为,结合互补模态对捕捉精细动作动态与行为上下文语义至关重要,可克服以往仅依赖动作或视频方法的局限。为此,我们提出ViMoNet,一种通过两阶段对齐与指令微调策略训练的多模态视觉-语言框架,结合精确的动作-文本监督与大规模视频-文本数据。我们还构建了包含动作序列、视频及指令级标注的VIMOS多模态数据集,以及用于评估行为中心推理的ViMoNet-Bench标准基准。实验表明,ViMoNet在描述生成、动作理解与行为解读任务中均显著优于现有方法。该框架在辅助医疗领域具有巨大潜力,如老年人监测、跌倒检测及老年健康风险早期识别。本工作有助于实现联合国可持续发展目标3(良好健康与福祉),推动可及性AI工具发展,促进全民健康覆盖,减少可预防健康问题,提升整体福祉。
原文摘要 · Abstract (English)
This study investigates the use of large language models (LLMs) for human behavior understanding by jointly leveraging motion and video data. We argue that integrating these complementary modalities is essential for capturing both fine-grained motion dynamics and contextual semantics of human actions, addressing the limitations of prior motion-only or video-only approaches. To this end, we propose ViMoNet, a multimodal vision-language framework trained through a two-stage alignment and instruction-tuning strategy that combines precise motion-text supervision with large-scale video-text data. We further introduce VIMOS, a multimodal dataset comprising human motion sequences, videos, and instruction-level annotations, along with ViMoNet-Bench, a standardized benchmark for evaluating behavior-centric reasoning. Experimental results demonstrate that ViMoNet consistently outperforms existing methods across caption generation, motion understanding, and human behavior interpretation tasks. The proposed framework shows significant potential in assistive healthcare applications, such as elderly monitoring, fall detection, and early identification of health risks in aging populations. This work contributes to the United Nations Sustainable Development Goal 3 (SDG 3: Good Health and Well-being) by enabling accessible AI-driven tools that promote universal health coverage, reduce preventable health issues, and enhance overall well-being.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。