arXiv:2506.16061cs.CV2025-06

用自适应超分辨率提升低清视频人体姿态估计精度与速度

STAR-Pose: Efficient Low-Resolution Video Human Pose Estimation via Spatial-Temporal Adaptive Super-Resolution

  • 设计时空自适应超分框架,融合改进的Transformer与并行CNN增强局部纹理
  • 在64x48低分辨率下实现5.2% mAP提升,推理速度比级联方法快2.8至4.4倍
  • 采用姿态感知损失,聚焦关键点定位所需结构特征,而非单纯追求画质

低分辨率视频中的人体姿态估计是计算机视觉中的基础挑战。传统方法要么假设输入为高质量图像,要么采用计算开销大的级联处理,限制了其在资源受限环境中的部署。本文提出STAR-Pose,一种专为视频人体姿态估计设计的时空自适应超分辨率框架。该方法引入一种带有LeakyReLU改进的线性注意力时空Transformer,有效捕捉长时序依赖;同时配备自适应融合模块,通过并行卷积网络增强局部纹理。我们还设计了一种姿态感知复合损失函数,引导网络重建对关键点定位最有益的结构特征,而非仅优化视觉质量。在多个主流视频人体姿态估计数据集上的大量实验表明,STAR-Pose表现优于现有方法:在极端低分辨率(64x48)条件下,mAP最高提升5.2%,推理速度较级联方法快2.8至4.4倍。

原文摘要 · Abstract (English)

Human pose estimation in low-resolution videos presents a fundamental challenge in computer vision. Conventional methods either assume high-quality inputs or employ computationally expensive cascaded processing, which limits their deployment in resource-constrained environments. We propose STAR-Pose, a spatial-temporal adaptive super-resolution framework specifically designed for video-based human pose estimation. Our method features a novel spatial-temporal Transformer with LeakyReLU-modified linear attention, which efficiently captures long-range temporal dependencies. Moreover, it is complemented by an adaptive fusion module that integrates parallel CNN branch for local texture enhancement. We also design a pose-aware compound loss to achieve task-oriented super-resolution. This loss guides the network to reconstruct structural features that are most beneficial for keypoint localization, rather than optimizing purely for visual quality. Extensive experiments on several mainstream video HPE datasets demonstrate that STAR-Pose outperforms existing approaches. It achieves up to 5.2% mAP improvement under extremely low-resolution (64x48) conditions while delivering 2.8x to 4.4x faster inference than cascaded approaches.

人体姿态估计超分辨率视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。