arXiv:2608.05782cs.CVcs.LG2026-08

用视频生成逼真的可控制惯性传感器数据,解决人体活动识别数据少的问题。

VSMP-IMU: Video-Grounded Semantic Motion Programs for Sensor-Aware Synthetic IMU Generation

论文配图:VSMP-IMU: Video-Grounded Semantic Motion Programs for Sensor-Aware Synthetic IMU Generation
图 1 · 摘自论文原文
  • 基于结构化动作语义程序,分离动作核心与变化细节。
  • 在五个数据集上平均宏准确率78.33%,比真实数据训练高9.77%。
  • 适合低资源、长尾分布场景,尤其适合数据稀缺的可穿戴设备研究者。

可穿戴人体活动识别(HAR)常受限于标注传感器数据稀少,尤其在低资源、类别不平衡和跨主体泛化场景下。合成惯性测量单元(IMU)数据可缓解此问题并提升模型性能,但现有方法存在权衡:视频驱动方法视觉对齐好但易受姿态估计误差影响,文本驱动方法可控性强但实际动作关联弱。本文提出VSMP-IMU,一种基于结构化语义动作程序(SMP)的视频引导合成框架,将动作定义语义与标签保持的变化分离开来。输入视频后,该框架提取并增强SMP,据此生成运动,转换为虚拟IMU信号,并将其适配至目标可穿戴领域。在五组公开的IMU-HAR数据集上,采用留一人的评估方式,VSMP-IMU平均宏F1达78.33%,相比仅使用真实数据训练提升9.77%,优于最强现有合成基线4.04%。在数据量减少的低资源设置下,相比真实数据训练提升18.54%,比最强合成基线平均提升超6%。在类别不平衡数据集的长尾评估中,尾部类别宏F1相比真实训练提升19.86%,相比最先进方法提升4.76%。结果表明,结构化的视频引导语义为可控、贴近可穿戴场景的合成传感器数据生成提供了有效基础。

原文摘要 · Abstract (English)

Wearable human activity recognition (HAR) is often limited by the scarcity of labeled sensor data, especially in low-resource, class-imbalanced, and subject-generalization settings. Synthetic IMU generation can reduce this dependency and enhance HAR machine learning model's performance, but existing approaches face a trade-off without addressing all factors: video-driven methods are visually grounded but sensitive to pose-estimation errors, while text-driven methods are controllable but often weakly grounded in how activities are actually performed. We present VSMP-IMU, a video-grounded framework for controllable synthetic IMU generation based on a structured Semantic Motion Program (SMP), which separates activity-defining semantics from label-preserving variation. Given an input video, VSMP-IMU extracts and augments an SMP, uses it to synthesize motion, converts the motion into virtual IMU signals, and grounds the resulting signals to the target wearable domain. We evaluate VSMP-IMU against state-of-the-art synthetic data generation methods on five public IMU-HAR datasets under leave-one-person-out evaluation. VSMP-IMU achieves an average Macro-F1 of 78.33%, improving over real-only training by 9.77% and over the strongest prior synthetic baseline by 4.04%. In low-resource settings with reduced training data-samples, it improves over real-only training by 18.54% and over the strongest prior synthetic baselines by more than 6% on average. Under long-tail evaluation in imbalanced datasets, it improves tail-class Macro-F1 by 19.86% over Real-only training and by 4.76% over SOTA. These results show that structured video-grounded semantics provide a practical foundation for controllable, wearable-relevant synthetic sensor data generation.

合成数据动作识别可穿戴设备视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。