用物理先验训练自监督模型,实现人体表面与骨骼的精细运动估计
H-Flow: Self-supervised Human Scene Flow via Physics-inspired Joint Multi-modal Learning

- 基于多模态联合学习,通过单目视频同时预测姿态与深度流
- 在无真实标注情况下,借助生物力学先验提升精度,零样本泛化至真实场景
- 构建高保真合成数据集DynAct4D,支持复杂服装与动作下的密集流标注
参数化人体模型可捕捉整体姿态,但无法表征衣物和软组织的非刚性形变;通用场景光流虽能估计密集运动,但在关节体上失效,且像素级标签难以获取。本文提出H-Flow,一种融合骨骼运动与表面变形的稠密人体场景光流方法。统一的多头变换器从单目视频中联合预测姿态与深度,作为辅助输出。由于缺乏标注,我们引入人体运动的物理先验,将几何、结构与生物力学约束作为跨模态训练目标。此外,我们构建了高保真合成基准DynAct4D,涵盖多样人物、服饰与动作的密集光流标注。在标准基准上,H-Flow优于场景光流与参数化基线,并实现零样本迁移至真实视频。代码、模型与动态4D数据集将在发表后开源。
原文摘要 · Abstract (English)
Parametric human models capture global pose but cannot represent the non-rigid surface dynamics of clothing and soft tissue. Generic scene flow estimates dense motion but breaks down on articulated bodies, where pixel-level supervision is also intractable to acquire. We introduce H-Flow, a dense human scene flow that captures both skeletal kinematics and surface deformation. A unified multi-head transformer estimates flow from monocular video, jointly predicting pose and depth as companion outputs. The challenge lies in the lack of supervision. In place of unattainable labels, we anchor the network in the physics of human motion, encoding geometric, structural, and biomechanical priors as cross-modal training objectives. We further introduce DynAct4D, a high-fidelity synthetic benchmark providing dense flow annotations across diverse subjects, garments, and motions. On standard benchmarks, H-Flow outperforms scene-flow and parametric baselines, and generalizes zero-shot to in-the-wild video. Code, models, and the DynAct4D benchmark will be released upon publication
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。