arXiv:2504.06397cs.CV2025-04CVPR被引 79

用提示词控制人体三维重建,兼顾场景上下文与精度。

PromptHMR: Promptable Human Mesh Recovery

  • 通过空间与语义提示重构人体姿态与形状
  • 在密集人群中小至人脸的框也能准确估计
  • 支持语言描述、交互标签等灵活控制

人体姿态与形状(HPS)估计在复杂场景如人群密集、人与人交互及单视角重建中面临挑战。现有方法缺乏融合辅助信息的机制,且高精度方法依赖裁剪的人体检测,无法利用场景上下文;而处理整图的方法常漏检且精度较低。尽管近期基于语言的方法尝试使用大语言或视觉-语言模型进行推理,但其指标表现仍远低于当前最优水平。为此,我们提出 PromptHMR,一种基于 Transformer 的可提示人体网格恢复方法,通过空间与语义提示重新建模 HPS 估计。该方法处理全图以保留场景上下文,支持多种输入模态:如边界框、掩码等空间提示,以及语言描述、交互标签等语义提示。实验表明,PromptHMR 在复杂场景下表现稳健:能从仅含人脸大小的边界框中估计人体,借助语言描述提升身体形状估计,建模人与人交互,并在视频中生成时序连贯动作。在多个基准测试上达到当前最优性能,同时提供灵活的提示式控制能力。

原文摘要 · Abstract (English)

Human pose and shape (HPS) estimation presents challenges in diverse scenarios such as crowded scenes, person-person interactions, and single-view reconstruction. Existing approaches lack mechanisms to incorporate auxiliary "side information" that could enhance reconstruction accuracy in such challenging scenarios. Furthermore, the most accurate methods rely on cropped person detections and cannot exploit scene context while methods that process the whole image often fail to detect people and are less accurate than methods that use crops. While recent language-based methods explore HPS reasoning through large language or vision-language models, their metric accuracy is well below the state of the art. In contrast, we present PromptHMR, a transformer-based promptable method that reformulates HPS estimation through spatial and semantic prompts. Our method processes full images to maintain scene context and accepts multiple input modalities: spatial prompts like bounding boxes and masks, and semantic prompts like language descriptions or interaction labels. PromptHMR demonstrates robust performance across challenging scenarios: estimating people from bounding boxes as small as faces in crowded scenes, improving body shape estimation through language descriptions, modeling person-person interactions, and producing temporally coherent motions in videos. Experiments on benchmarks show that PromptHMR achieves state-of-the-art performance while offering flexible prompt-based control over the HPS estimation process.

人体重建提示工程多模态视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。