用视觉语言模型引导,提升扩散模型人体建模的准确与合理。
VLM-Guided Group Preference Alignment for Diffusion-based Human Mesh Recovery
- 引入双重记忆自省式评估器,生成上下文感知的质量评分
- 构建群体偏好数据集,使模型生成更符合物理规律且对齐图像的网格
- 在遮挡和复杂场景下表现更优,适合真实世界人体重建任务
从单张RGB图像恢复人体网格(HMR)本质上存在歧义,同一2D观察可能对应多个3D姿态。近期基于扩散的方法通过生成多种假设来应对,但常牺牲准确性,导致预测结果在遮挡或复杂户外场景中出现物理不合理或偏离输入图像的问题。为此,我们提出一种双记忆增强型HMR评述代理,具备自我反思能力,可生成上下文感知的质量评分,提炼关于3D人体运动结构、物理可行性及与输入图像对齐程度的细粒度线索。利用这些评分构建群体偏好数据集,并提出群体偏好对齐框架用于微调扩散式HMR模型。该过程将丰富的偏好信号注入模型,引导其生成更物理合理且与图像一致的人体网格。大量实验表明,本方法优于现有最先进方法。
原文摘要 · Abstract (English)
Human mesh recovery (HMR) from a single RGB image is inherently ambiguous, as multiple 3D poses can correspond to the same 2D observation. Recent diffusion-based methods tackle this by generating various hypotheses, but often sacrifice accuracy. They yield predictions that are either physically implausible or drift from the input image, especially under occlusion or in cluttered, in-the-wild scenes. To address this, we introduce a dual-memory augmented HMR critique agent with self-reflection to produce context-aware quality scores for predicted meshes. These scores distill fine-grained cues about 3D human motion structure, physical feasibility, and alignment with the input image. We use these scores to build a group-wise HMR preference dataset. Leveraging this dataset, we propose a group preference alignment framework for finetuning diffusion-based HMR models. This process injects the rich preference signals into the model, guiding it to generate more physically plausible and image-consistent human meshes. Extensive experiments demonstrate that our method achieves superior performance compared to state-of-the-art approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。