arXiv:2505.20920cs.CV2025-05CVPR被引 5

通过融合视频与动作特征,提升人体运动理解中的概念发现能力。

HuMoCon: Concept Discovery for Human Motion Understanding

  • 用视频上下文对齐动作特征,增强细粒度交互建模。
  • 引入速度重建机制,保留高频运动信息,缓解时间过平滑。
  • 适用于大模型训练,适合行为分析与动作理解研究者使用。

我们提出HuMoCon,一种用于高级人体行为分析的新型运动-视频理解框架。其核心是人类运动概念发现框架,可高效训练多模态编码器以提取语义有意义且泛化性强的特征。该方法解决了运动概念发现中缺乏显式多模态特征对齐、掩码自编码框架中高频信息丢失等关键问题。通过利用视频提供上下文理解、运动实现细粒度交互建模,并引入速度重建机制以增强高频特征表达并缓解时间过平滑。在标准基准上的全面实验表明,HuMoCon能有效实现运动概念发现,在训练大型模型进行人体运动理解方面显著优于现有先进方法。论文将开源相关代码。

原文摘要 · Abstract (English)

We present HuMoCon, a novel motion-video understanding framework designed for advanced human behavior analysis. The core of our method is a human motion concept discovery framework that efficiently trains multi-modal encoders to extract semantically meaningful and generalizable features. HuMoCon addresses key challenges in motion concept discovery for understanding and reasoning, including the lack of explicit multi-modality feature alignment and the loss of high-frequency information in masked autoencoding frameworks. Our approach integrates a feature alignment strategy that leverages video for contextual understanding and motion for fine-grained interaction modeling, further with a velocity reconstruction mechanism to enhance high-frequency feature expression and mitigate temporal over-smoothing. Comprehensive experiments on standard benchmarks demonstrate that HuMoCon enables effective motion concept discovery and significantly outperforms state-of-the-art methods in training large models for human motion understanding. We will open-source the associated code with our paper.

运动理解多模态概念发现视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。