用文字或位置提示精准定位图像中的人体姿态与掩码
Referring Human Pose and Mask Estimation in the Wild
- 提出统一可提示的端到端模型UniPHD,融合多模态信息
- 在5万+野外场景实例上实现高精度身份感知的姿态与掩码预测
- 适合需要精准人体理解的机器人辅助、运动分析等场景
我们提出了野外环境中的指代人体姿态与掩码估计(R-HPM)任务,用户可通过文本或位置提示指定图像中关注的人。该任务在助人机器人、运动分析等以人为中心的应用中具有重要意义。与以往工作不同,R-HPM同时实现高质量、身份感知的姿态与掩码输出。为此,我们构建了大规模数据集RefHuman,基于MS COCO扩展超过50,000个野外场景实例,包含关键点、掩码及文本/位置提示标注。为支持提示驱动的估计,我们提出首个端到端可提示方法UniPHD,通过多模态特征提取和姿态中心的分层解码器,处理文本或位置实例查询与关键点查询,生成针对所指个体的结果。大量实验表明,UniPHD在用户友好提示下表现优异,在RefHuman验证集与MS COCO val2017上均达到顶尖性能。
原文摘要 · Abstract (English)
We introduce Referring Human Pose and Mask Estimation (R-HPM) in the wild, where either a text or positional prompt specifies the person of interest in an image. This new task holds significant potential for human-centric applications such as assistive robotics and sports analysis. In contrast to previous works, R-HPM (i) ensures high-quality, identity-aware results corresponding to the referred person, and (ii) simultaneously predicts human pose and mask for a comprehensive representation. To achieve this, we introduce a large-scale dataset named RefHuman, which substantially extends the MS COCO dataset with additional text and positional prompt annotations. RefHuman includes over 50,000 annotated instances in the wild, each equipped with keypoint, mask, and prompt annotations. To enable prompt-conditioned estimation, we propose the first end-to-end promptable approach named UniPHD for R-HPM. UniPHD extracts multimodal representations and employs a proposed pose-centric hierarchical decoder to process (text or positional) instance queries and keypoint queries, producing results specific to the referred person. Extensive experiments demonstrate that UniPHD produces quality results based on user-friendly prompts and achieves top-tier performance on RefHuman val and MS COCO val2017. Data and Code: https://github.com/bo-miao/RefHuman
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。