arXiv:2504.20800cs.CV2025-04被引 2

用RGB图像频域特征提升人体感知预训练效果,无需深度图数据

Adept: Annotation-Denoising Auxiliary Tasks with Discrete Cosine Transform Map and Keypoint for Human-Centric Pretraining

  • 通过DCT将RGB图像转至频域,挖掘深层语义信息
  • 引入关键点与DCT图作为去噪辅助任务,提升模型精度
  • 适用于人体姿态、分割、计数等多任务,尤其适合无深度数据场景

人体中心感知是多种计算机视觉任务的核心,但以往研究多针对单一任务,性能受限于特定数据集规模。近期方法虽利用深度信息学习细粒度语义,却受相机视角敏感性及互联网上稀少的RGB-D数据制约。本文舍弃深度信息,通过离散余弦变换(DCT)在频域中探索RGB图像的语义特征,并提出基于关键点和DCT图的标注去噪辅助任务,强化模型对人身体细粒度语义的学习能力。大量实验表明,在未使用深度标注的大规模数据集(COCO和AIC)上预训练时,模型在人体姿态估计上比当前最优方法提升+0.5 mAP(COCO)、+1.4 PCKh(MPII)、-0.51 EPE(Human3.6M);在人体解析上提升+4.50 mIoU(Human3.6M);在人群计数上降低-3.14 MAE(SHA)、-0.07 MAE(SHB);在人群定位上提升+1.1 F1(SHA)、+0.8 F1(SHB);在行人重识别上提升+0.1 mAP(Market1501)、+0.8 mAP(MSMT)。同时在MPII+NTURGBD数据集上验证了方法有效性。

原文摘要 · Abstract (English)

Human-centric perception is the core of diverse computer vision tasks and has been a long-standing research focus. However, previous research studied these human-centric tasks individually, whose performance is largely limited to the size of the public task-specific datasets. Recent human-centric methods leverage the additional modalities, e.g., depth, to learn fine-grained semantic information, which limits the benefit of pretraining models due to their sensitivity to camera views and the scarcity of RGB-D data on the Internet. This paper improves the data scalability of human-centric pretraining methods by discarding depth information and exploring semantic information of RGB images in the frequency space by Discrete Cosine Transform (DCT). We further propose new annotation denoising auxiliary tasks with keypoints and DCT maps to enforce the RGB image extractor to learn fine-grained semantic information of human bodies. Our extensive experiments show that when pretrained on large-scale datasets (COCO and AIC datasets) without depth annotation, our model achieves better performance than state-of-the-art methods by +0.5 mAP on COCO, +1.4 PCKh on MPII and -0.51 EPE on Human3.6M for pose estimation, by +4.50 mIoU on Human3.6M for human parsing, by -3.14 MAE on SHA and -0.07 MAE on SHB for crowd counting, by +1.1 F1 score on SHA and +0.8 F1 score on SHA for crowd localization, and by +0.1 mAP on Market1501 and +0.8 mAP on MSMT for person ReID. We also validate the effectiveness of our method on MPII+NTURGBD datasets

人体感知预训练无监督频域建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。