arXiv:2605.14854cs.CVcs.AI2026-05

分治式人体三维重建,精准恢复躯干,灵活处理四肢模糊。

FactorizedHMR: A Hybrid Framework for Video Human Mesh Recovery

论文配图:FactorizedHMR: A Hybrid Framework for Video Human Mesh Recovery
图 1 · 摘自论文原文
  • 先用确定性模块稳住躯干和根部,再用概率流匹配补全四肢
  • 在遮挡场景下误差降低17.3%,世界坐标系追踪更稳定
  • 适合需要高精度肢体姿态的虚拟人、动作捕捉应用

人体网格恢复(HMR)本质上存在歧义:在遮挡或弱深度线索下,多个三维人体结构可能解释同一图像。这种歧义在身体各部位分布不均——躯干和根部姿态通常较受约束,而四肢等远端关节更不确定。基于此,我们提出FactorizedHMR,一种两阶段框架,对不同区域采用差异化处理。首先通过确定性回归模块恢复稳定的躯干-根部锚点;随后利用概率流匹配模块完成其余非躯干关节的重建。为提升重建可靠性,结合复合目标表示、几何感知监督与特征感知无分类器引导,在保持躯干锚点不变的同时,显著改善了易模糊关节的单参考恢复效果。我们还设计了一条合成数据流水线,提供多样视角下的图像-相机-运动配对监督。在相机空间与世界空间基准测试中,FactorizedHMR性能优于强基线,尤其在遮挡密集场景和对漂移敏感的世界空间指标上表现突出。

原文摘要 · Abstract (English)

Human Mesh Recovery (HMR) is fundamentally ambiguous: under occlusion or weak depth cues, multiple 3D bodies can explain the same image evidence. This ambiguity is not uniform across the body, as torso pose and root structure are often relatively well constrained, whereas distal articulations such as the arms and legs are more uncertain. Building on this observation, we propose FactorizedHMR, a two-stage framework that treats these two regimes differently. A deterministic regression module first recovers a stable torso-root anchor, and a probabilistic flow-matching module then completes the remaining non-torso articulation. To make this completion reliable, we combine a composite target representation with geometry-aware supervision and feature-aware classifier-free guidance, preserving the torso-root anchor while improving single-reference recovery of ambiguity-prone articulation. We also introduce a synthetic data pipeline that provides the paired image-camera-motion supervision under diverse viewpoints. Across camera-space and world-space benchmarks, FactorizedHMR remains competitive with strong baselines, with the clearest gains in occlusion-heavy recovery and drift-sensitive world-space metrics.

人体重建视频恢复概率建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。