arXiv:2604.21712cs.CVcs.MM2026-04

融合视觉与生成优势,提升遮挡下3D人体建模精度

Discriminative-Generative Synergy for Occlusion Robust 3D Human Mesh Recovery

论文配图:Discriminative-Generative Synergy for Occlusion Robust 3D Human Mesh Recovery
图 1 · 摘自论文原文
  • 用视觉变压器提取可见区域特征,扩散模型生成遮挡部分
  • 在多种遮挡场景下,关键指标优于现有方法
  • 适合需要高鲁棒性3D人体重建的工业应用

从单目RGB图像恢复3D人体网格旨在为下游应用生成解剖学合理的3D人体模型,但在部分或严重遮挡下仍具挑战。基于回归的方法效率高但常在非约束场景中产生不合理的结果;基于扩散的方法虽能提供遮挡区域的强生成先验,但可能因过度依赖生成而削弱对罕见姿态的保真度。为此,我们提出一种受大脑启发的协同框架,将视觉变压器的判别能力与条件扩散模型的生成能力相结合。具体地,基于ViT的路径从可见区域提取确定性视觉线索,而扩散路径则合成结构一致的人体表示。为有效连接两条路径,我们设计了多样-一致特征学习模块以对齐判别特征与生成先验,并引入跨注意力多层级融合机制,实现语义层级间的双向交互。在标准基准上的实验表明,该方法在关键指标上表现更优,且在复杂真实场景中展现出强鲁棒性。

原文摘要 · Abstract (English)

3D human mesh recovery from monocular RGB images aims to estimate anatomically plausible 3D human models for downstream applications, but remains challenging under partial or severe occlusions. Regression-based methods are efficient yet often produce implausible or inaccurate results in unconstrained scenarios, while diffusion-based methods provide strong generative priors for occluded regions but may weaken fidelity to rare poses due to over-reliance on generation. To address these limitations, we propose a brain-inspired synergistic framework that integrates the discriminative power of vision transformers with the generative capability of conditional diffusion models. Specifically, the ViT-based pathway extracts deterministic visual cues from visible regions, while the diffusion-based pathway synthesizes structurally coherent human body representations. To effectively bridge the two pathways, we design a diverse-consistent feature learning module to align discriminative features with generative priors, and a cross-attention multi-level fusion mechanism to enable bidirectional interaction across semantic levels. Experiments on standard benchmarks demonstrate that our method achieves superior performance on key metrics and shows strong robustness in complex real-world scenarios.

3D人体重建扩散模型遮挡鲁棒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。