arXiv:2503.11194cs.CV2025-03被引 1

解决3D人体姿态估计在真实视频流中使用估算2D姿态的自适应难题

Online Test-time Adaptation for 3D Human Pose Estimation: A Practical Perspective with Estimated 2D Poses

  • 通过自适应聚合初始化模型,分阶段优化减少错误更新影响
  • 引入局部增强,用邻近可信样本修正当前不可靠预测
  • 实测在真实估算2D姿态场景下显著超越现有方法

在线测试时自适应用于处理与训练数据不同的视频流中的3D人体姿态估计。传统方法依赖真实2D姿态进行适应,但实际中仅能获取估计的2D姿态。本文针对使用估计2D姿态的视频流,提出新方法。分析表明,在限制估计误差的同时保留准确姿态信息是核心挑战。为此,提出自适应聚合、两阶段优化和局部增强策略:首先跨视频进行自适应聚合,以标注样本初始化模型状态;其次在每段视频内采用两阶段优化,在利用2D拟合优势的同时最小化错误更新的影响;最后通过局部增强,利用相邻高置信度样本更新模型,再适应当前低置信度样本。实验表明,该方法在估计2D姿态条件下显著优于现有最先进方法,推动了实际应用场景下的自适应发展。

原文摘要 · Abstract (English)

Online test-time adaptation for 3D human pose estimation is used for video streams that differ from training data. Ground truth 2D poses are used for adaptation, but only estimated 2D poses are available in practice. This paper addresses adapting models to streaming videos with estimated 2D poses. Comparing adaptations reveals the challenge of limiting estimation errors while preserving accurate pose information. To this end, we propose adaptive aggregation, a two-stage optimization, and local augmentation for handling varying levels of estimated pose error. First, we perform adaptive aggregation across videos to initialize the model state with labeled representative samples. Within each video, we use a two-stage optimization to benefit from 2D fitting while minimizing the impact of erroneous updates. Second, we employ local augmentation, using adjacent confident samples to update the model before adapting to the current non-confident sample. Our method surpasses state-of-the-art by a large margin, advancing adaptation towards more practical settings of using estimated 2D poses.

3D姿态估计在线自适应视频流姿态估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。