arXiv:2409.06535cs.CV2024-09ECCV被引 12

用多模态融合提升3D人体姿态表示,支持任意组合输入

PoseEmbroider: Towards a 3D, Visual, Semantic-aware Human Pose Representation

  • 基于Transformer的检索式模型,融合图像、3D姿态与文本描述
  • 在部分信息缺失时仍能准确推断姿态,优于传统多模态对齐模型
  • 适用于姿态估计与精细动作指令生成,无需重新训练

将多模态信息(如图像与文本)对齐至潜在空间,已显著提升语义视觉表示能力,推动图像字幕、文本到图像生成等任务。在人体视觉领域,尽管类似CLIP的表示可较好编码常见姿势(如站立或坐姿),但对细节或非标准姿势的区分能力不足。虽然3D人体姿态常与图像(如姿态估计或姿态条件生成)或文本(如文本到姿态生成)关联,却很少同时结合三者。本文提出一种新方法,融合3D姿态、人物图像与文本描述,构建更精准的3D、视觉与语义感知的人体姿态表示。引入基于Transformer的模型,采用检索式训练方式,可接受上述任一组合的输入。在多模态组合时,性能超越标准多模态对齐检索模型,支持部分信息缺失场景(如下半身遮挡)。该表示可用于:(1) 从图像中回归SMPL参数,可选地加入文本提示;(2) 细粒度指令生成任务,即生成从一个3D姿态移动到另一个的姿态指导说明(如健身教练)。不同于以往工作,本模型可处理任意输入组合,无需重新训练。

原文摘要 · Abstract (English)

Aligning multiple modalities in a latent space, such as images and texts, has shown to produce powerful semantic visual representations, fueling tasks like image captioning, text-to-image generation, or image grounding. In the context of human-centric vision, albeit CLIP-like representations encode most standard human poses relatively well (such as standing or sitting), they lack sufficient acuteness to discern detailed or uncommon ones. Actually, while 3D human poses have been often associated with images (e.g. to perform pose estimation or pose-conditioned image generation), or more recently with text (e.g. for text-to-pose generation), they have seldom been paired with both. In this work, we combine 3D poses, person's pictures and textual pose descriptions to produce an enhanced 3D-, visual- and semantic-aware human pose representation. We introduce a new transformer-based model, trained in a retrieval fashion, which can take as input any combination of the aforementioned modalities. When composing modalities, it outperforms a standard multi-modal alignment retrieval model, making it possible to sort out partial information (e.g. image with the lower body occluded). We showcase the potential of such an embroidered pose representation for (1) SMPL regression from image with optional text cue; and (2) on the task of fine-grained instruction generation, which consists in generating a text that describes how to move from one 3D pose to another (as a fitness coach). Unlike prior works, our model can take any kind of input (image and/or pose) without retraining.

3D姿态多模态生成模型姿势理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。