首个融合肌电与多视角的健身动作质量评估数据集,支持精准动作分析。
FLEX: A Largescale Multimodal, Multiview Dataset for Learning Structured Representations for Fitness Action Quality Assessment
- 采集38人完成20种负重动作的多视角视频、3D姿态与肌电信号
- 包含超7500段数据,专家标注构建可解释的健身知识图谱
- 支持跨模态推理与结构化问答,适合智能健身教练研究
动作质量评估(AQA)在健身房力量训练中具有重要应用价值,能有效预防损伤并提升训练效果。现有数据集局限于单视角竞技体育和RGB视频,缺乏多模态信号与专业评估。本文提出FLEX,首个大规模多模态、多视角健身AQA数据集,集成表面肌电(sEMG)。FLEX包含38名不同技能水平受试者完成20种负重动作的超过7,500段多视角记录,同步采集RGB视频、3D姿态、sEMG及生理信号。专家标注形成健身知识图谱(FKG),关联动作、关键步骤、错误类型与反馈,支持可解释的组合评分。FLEX支持多模态融合、跨模态预测(包括新颖的Video→EMG任务)及生物力学导向表示学习。基于FKG,进一步构建FLEX-VideoQA结构化问答基准,实现分层查询驱动的视觉-语言模型跨模态推理。基线实验表明,多模态输入、多视角视频与细粒度标注显著提升AQA性能。FLEX推动AQA向更丰富的多模态场景发展,为人工智能健身评估与教练系统提供基础。数据与代码已开源。
原文摘要 · Abstract (English)
Action Quality Assessment (AQA) -- the task of quantifying how well an action is performed -- has great potential for detecting errors in gym weight training, where accurate feedback is critical to prevent injuries and maximize gains. Existing AQA datasets, however, are limited to single-view competitive sports and RGB video, lacking multimodal signals and professional assessment of fitness actions. We introduce FLEX, the first large-scale, multimodal, multiview dataset for fitness AQA that incorporates surface electromyography (sEMG). FLEX contains over 7,500 multiview recordings of 20 weight-loaded exercises performed by 38 subjects of diverse skill levels, with synchronized RGB video, 3D pose, sEMG, and physiological signals. Expert annotations are organized into a Fitness Knowledge Graph (FKG) linking actions, key steps, error types, and feedback, supporting a compositional scoring function for interpretable quality assessment. FLEX enables multimodal fusion, cross-modal prediction -- including the novel Video$\rightarrow$EMG task -- and biomechanically oriented representation learning. Building on the FKG, we further introduce FLEX-VideoQA, a structured question-answering benchmark with hierarchical queries that drive cross-modal reasoning in vision-language models. Baseline experiments demonstrate that multimodal inputs, multiview video, and fine-grained annotations significantly enhance AQA performance. FLEX thus advances AQA toward richer multimodal settings and provides a foundation for AI-powered fitness assessment and coaching. Dataset and code are available at \href{https://github.com/HaoYin116/FLEX}{https://github.com/HaoYin116/FLEX}. Link to Project \href{https://haoyin116.github.io/FLEX_Dataset}{page}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。