arXiv:2609.01840cs.CV2026-09

用婴儿视频蒸馏模型,让3D姿态估计更准。

Cross-Model Distillation of a Human-Pose Foundation Model from Unannotated Infant Video for Markerless 3D Pose Estimation

论文配图:Cross-Model Distillation of a Human-Pose Foundation Model from Unannotated Infant Video for Markerless 3D Pose Estimation
图 1 · 摘自论文原文
  • 用成人模型生成伪标签,指导婴儿专用模型优化
  • 2D关键点准确率提升至42%,3D位置误差降至22.2mm
  • 无需人工标注,适合临床早期筛查应用

自发运动是评估婴儿神经运动健康的重要窗口,现有临床评分工具虽能预测脑瘫风险,但依赖专业人员、耗时且存在评分者差异。这推动了基于视频的无标记自动评估发展,而标记捕捉在婴儿中不现实。然而,当前用于无标记捕捉的基础模型几乎全部在成人数据上训练:我们之前的多视角婴儿研究发现,没有一个模型在2D关键点精度和直接3D身体重建两方面同时最优,性能存在明显权衡。本研究通过跨模型蒸馏,将Sapiens 2姿态模型的知识迁移至SAM 3D Body模型,仅使用未标注婴儿视频。冻结的教师模型提供密集伪标签,可微渲染器在训练中对齐预测网格与伪标签。在11名婴儿(18个会话,173段视频)上,按先前研究的多视角协议测试,微调后同视角2D关键点正确率(PCK@10px)从0.22提升至0.42(身体与面部),经普罗克鲁斯特斯对齐的平均关节3D位置误差从25.5mm降至22.2mm。结果表明,跨模型蒸馏有效提升了SAM 3D Body模型在婴儿上的表现。

原文摘要 · Abstract (English)

Spontaneous movement is one of the earliest windows onto an infant's neuromotor health, and structured clinical instruments that score it are validated early predictors of cerebral-palsy risk. However, they require specially trained raters, are time-consuming, and carry inter-rater variability. This motivates automated, video-based markerless assessment, especially as marker-based motion capture is impractical in infants. Yet the foundation models that make markerless capture possible are trained almost entirely on adults: our recent multi-view infant study found that no single model is jointly best, with strong 2D keypoint accuracy and direct 3D body recovery split across different models. While that study identifies this trade-off, it does not resolve it. Here, we perform cross-model distillation from the Sapiens 2 pose model into the SAM 3D Body model, using unannotated infant video alone. A frozen teacher supplies dense pseudo-labels, and a differentiable renderer aligns the predicted mesh to them in the training loop. On eleven held-out infants (18 sessions, 173 recordings) under our prior study's multi-view protocol, fine-tuning improves same-view 2D keypoint agreement with the Sapiens reference (median body percentage of correct keypoints @ 10px 0.22 -> 0.42, face 0.22 -> 0.42) and Procrustes-aligned mean per joint 3D position error (25.5 -> 22.2 mm). This demonstrates how cross-model distillation improves SAM 3D Body model performance on infants.

姿态估计婴儿识别模型蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。