arXiv:2502.03836cs.CV2025-02被引 1

用视觉语言模型提升单目人体网格重建的精度与一致性

Adapting Human Mesh Recovery with Vision-Language Feedback

  • 结合2D图像与文本描述,通过对比学习对齐文本与姿态表示
  • 利用扩散模型优化初始参数,显著提升3D感知与图像对齐效果
  • 适合关注多模态融合与人体建模的研究者

人体网格恢复可采用基于回归或基于优化的方法。回归模型虽姿态准确,但因缺乏显式的2D-3D对应关系,难以实现模型与图像对齐;优化方法虽能对齐3D模型与2D观测,却易陷入局部极小值且存在深度歧义。本文利用大规模视觉语言模型(VLM)生成交互式身体部位描述,作为隐式约束以增强3D感知并缩小优化空间。具体地,将单目人体网格恢复建模为分布适应任务,融合2D观测与语言描述。为弥合文本与3D姿态信号之间的差距,首先训练文本编码器与姿态VQ-VAE,通过对比学习在共享潜在空间中对齐文本与身体姿态。随后,采用基于扩散的框架,依据来自2D观测与文本描述的梯度,对初始参数进行精炼。最终模型可生成具有精确3D感知与图像一致性的姿态。多个基准测试结果验证了其有效性。代码将公开。

原文摘要 · Abstract (English)

Human mesh recovery can be approached using either regression-based or optimization-based methods. Regression models achieve high pose accuracy but struggle with model-to-image alignment due to the lack of explicit 2D-3D correspondences. In contrast, optimization-based methods align 3D models to 2D observations but are prone to local minima and depth ambiguity. In this work, we leverage large vision-language models (VLMs) to generate interactive body part descriptions, which serve as implicit constraints to enhance 3D perception and limit the optimization space. Specifically, we formulate monocular human mesh recovery as a distribution adaptation task by integrating both 2D observations and language descriptions. To bridge the gap between text and 3D pose signals, we first train a text encoder and a pose VQ-VAE, aligning texts to body poses in a shared latent space using contrastive learning. Subsequently, we employ a diffusion-based framework to refine the initial parameters guided by gradients derived from both 2D observations and text descriptions. Finally, the model can produce poses with accurate 3D perception and image consistency. Experimental results on multiple benchmarks validate its effectiveness. The code will be made publicly available.

人体建模视觉语言模型扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。