arXiv:2409.07798cs.CV2024-09

用注意力与门控卷积提升姿态估计,更准更快。

GateAttentionPose: Enhancing Pose Estimation with Agent Attention and Improved Gated Convolutions

  • 用代理注意力替代大核卷积,兼顾全局信息与效率
  • 在COCO和MPII数据集上超越现有方法,精度更高
  • 适合自动驾驶、虚拟现实等复杂场景应用

本文提出GateAttentionPose,改进UniRepLKNet用于姿态估计。核心贡献包括:代理注意力模块(Agent Attention)取代大核卷积,在保持全局上下文建模能力的同时显著提升计算效率;门控增强前馈模块(GEFB)增强特征提取与处理能力,尤其在复杂场景中表现优异。在COCO和MPII数据集上的大量实验表明,该方法优于现有先进模型,包括原始UniRepLKNet,实现了更高精度或相当精度的同时具备更优效率。本方法为自动驾驶、人体动作捕捉及虚拟现实等多样化应用提供了稳健的姿态估计解决方案。

原文摘要 · Abstract (English)

This paper introduces GateAttentionPose, an innovative approach that enhances the UniRepLKNet architecture for pose estimation tasks. We present two key contributions: the Agent Attention module and the Gate-Enhanced Feedforward Block (GEFB). The Agent Attention module replaces large kernel convolutions, significantly improving computational efficiency while preserving global context modeling. The GEFB augments feature extraction and processing capabilities, particularly in complex scenes. Extensive evaluations on COCO and MPII datasets demonstrate that GateAttentionPose outperforms existing state-of-the-art methods, including the original UniRepLKNet, achieving superior or comparable results with improved efficiency. Our approach offers a robust solution for pose estimation across diverse applications, including autonomous driving, human motion capture, and virtual reality.

姿态估计注意力机制卷积网络高效模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。