构建首个大规模多模态动作理解数据集,支持细粒度分析与因果推理。
A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning
- 用提示工程生成逻辑连贯的动作描述,提升标注一致性。
- 覆盖40类动作、58,445样本,平均识别准确率76.52%。
- 适合研究多模态行为理解、大模型推理的学者和工程师。
多模态人体动作识别(HAR)利用互补传感器进行动作分类。除了识别外,大语言模型(LLMs)的发展使详细描述与因果推理成为可能,催生了人体动作理解(HAU)和推理(HARn)新任务。然而,现有大模型尤其是大视觉语言模型(LVLMs)在深度、IMU、毫米波等非RGB模态上表现不佳,主要因缺乏大规模数据-文本资源。现有HAR数据集仅提供粗粒度标签,难以捕捉精细动作动态。本文提出两种真实标注类型:(1)数据标签(离散类别),(2)数据描述(文本)。直接从标签生成描述易出现逻辑与时空不一致。为此,我们引入CUHK-X,一个大规模多模态数据集与基准套件,用于HAR、HAU与HARn。CUHK-X包含58,445个样本,涵盖30名参与者在两个室内环境执行的40类动作。为提升描述一致性,我们提出基于提示的场景生成方法,利用LLM生成逻辑连贯的动作序列,并经人工验证。该数据集包含三个基准,共六个评估任务。实验显示平均准确率为76.52%(HAR)、40.76%(HAU)、70.25%(HARn)。CUHK-X旨在推动社区发展数据密集型学习方法,实现鲁棒的多模态人体活动分析。项目页面与代码:https://openaiotlab.github.io/CUHK-X/ 与 https://github.com/openaiotlab/CUHK-X。
原文摘要 · Abstract (English)
Multimodal human action recognition (HAR) leverages complementary sensors for activity classification. Beyond recognition, recent advances in large language models (LLMs) enable detailed descriptions and causal reasoning, motivating new tasks: human action understanding (HAU) and human action reasoning (HARn). However, most LLMs, especially large vision language models (LVLMs), struggle with non-RGB modalities such as depth, IMU, and mmWave due to the lack of large-scale data-caption resources. Existing HAR datasets mainly provide coarse data-label annotations, which are insufficient to capture fine-grained action dynamics needed for HAU and HARn. We consider two ground-truth pair types: (1) data label (discrete category) and (2) data caption (textual description). Naively generating captions from labels often lacks logical and spatiotemporal consistency. We introduce CUHK-X, a large-scale multimodal dataset and benchmark suite for HAR, HAU, and HARn. CUHK-X contains 58,445 samples covering 40 actions performed by 30 participants across two indoor environments. To improve caption consistency, we propose a prompt-based scene creation method that leverages LLMs to generate logically connected activity sequences, followed by human validation. CUHK-X includes three benchmarks with six evaluation tasks. Experiments report average accuracies of 76.52% (HAR), 40.76% (HAU), and 70.25% (HARn). CUHK-X aims to enable the community to apply and develop data-intensive learning methods for robust, multimodal human activity analysis. Project page and code: https://openaiotlab.github.io/CUHK-X/ and https://github.com/openaiotlab/CUHK-X.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。