arXiv:2504.19032cs.CVcs.AI2025-04被引 2

用动态中心点统一人体姿态与分割,提升复杂场景下的准确性和速度

VISUALCENT: Visual Human Analysis using Dynamic Centroid Representation

  • 基于中心点的自底向上检测,结合盘状热图定位关键点
  • 在COCO和OCHuman上实现更高mAP与每秒超30帧的实时性能
  • 适合需要高精度多人分析的实时视觉系统应用

我们提出VISUALCENT,一种统一的人体姿态与实例分割框架,以解决多人群体视觉人体分析中的泛化性与可扩展性问题。该方法采用基于中心点的自底向上关键点检测范式,通过融合盘状表示(Disk Representation)与关键中心点(KeyCentroid)的热图来确定最优关键点坐标。针对统一分割任务,定义显式关键点为动态中心点(MaskCentroid),在人体快速运动或严重遮挡环境下可迅速将像素聚类至对应人体实例。在COCO和OCHuman数据集上的实验表明,VISUALCENT在mAP指标和每秒处理帧数(FPS)上均优于现有方法。代码已在项目页面开源。

原文摘要 · Abstract (English)

We introduce VISUALCENT, a unified human pose and instance segmentation framework to address generalizability and scalability limitations to multi person visual human analysis. VISUALCENT leverages centroid based bottom up keypoint detection paradigm and uses Keypoint Heatmap incorporating Disk Representation and KeyCentroid to identify the optimal keypoint coordinates. For the unified segmentation task, an explicit keypoint is defined as a dynamic centroid called MaskCentroid to swiftly cluster pixels to specific human instance during rapid changes in human body movement or significantly occluded environment. Experimental results on COCO and OCHuman datasets demonstrate VISUALCENTs accuracy and real time performance advantages, outperforming existing methods in mAP scores and execution frame rate per second. The implementation is available on the project page.

人体姿态估计实例分割动态中心点实时分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。