提出瀑布式Transformer结构,提升多人姿态估计精度
Waterfall Transformer for Multi-person Pose Estimation
- 采用分层瀑布模块生成多尺度特征图
- 在COCO数据集上优于其他Transformer架构
- 适合需要高精度姿态识别的实时应用
我们提出WTPose框架,一种用于多人姿态估计的单阶段、端到端可训练架构。该框架基于Transformer的瀑布模块,从不同骨干网络阶段生成多尺度特征图。模块在级联结构中执行滤波操作,扩展感受野并捕捉局部与全局上下文,从而增强网络整体特征表示能力。在COCO数据集上的实验表明,采用改进Swin骨干和基于Transformer的瀑布模块的WTPose架构,在多人姿态估计任务中优于其他Transformer架构。
原文摘要 · Abstract (English)
We propose the Waterfall Transformer architecture for Pose estimation (WTPose), a single-pass, end-to-end trainable framework designed for multi-person pose estimation. Our framework leverages a transformer-based waterfall module that generates multi-scale feature maps from various backbone stages. The module performs filtering in the cascade architecture to expand the receptive fields and to capture local and global context, therefore increasing the overall feature representation capability of the network. Our experiments on the COCO dataset demonstrate that the proposed WTPose architecture, with a modified Swin backbone and transformer-based waterfall module, outperforms other transformer architectures for multi-person pose estimation
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。