arXiv:2601.02267cs.CV2026-01被引 1

用扩散模型生成稠密对应点,提升多视角人体建模精度

DiffProxy: Multi-View Human Mesh Recovery via Diffusion-Generated Dense Proxies

  • 用预训练扩散模型生成像素到表面的稠密对应关系
  • 在五个真实数据集上达到当前最优性能,仅用合成数据训练
  • 适合需要高精度人体重建的研究者和工业应用

从多视角图像中精确恢复人体网格仍具挑战:端到端方法产生难以定位的纠缠误差,而基于拟合的方法依赖稀疏关键点,表面约束有限。我们发现瓶颈在于中间表示质量,且可通过复用具有丰富视觉先验的预训练扩散模型生成稠密像素-表面对应关系。提出DiffProxy,一个基于Stable Diffusion的框架,在大规模合成数据上训练,带有像素级标注。多条件代理生成器从多视角图像预测稠密对应关系,提供均匀表面约束以实现精准拟合。手部细化模块结合放大手部区域与全身图像进行细粒度优化,测试时缩放利用扩散模型随机性估计像素级不确定性。仅在合成数据上训练,但在五个不同真实世界基准上取得当前最佳结果。

原文摘要 · Abstract (English)

Precise human mesh recovery (HMR) from multi-view images remains challenging: end-to-end methods produce entangled errors hard to localize, while fitting-based methods rely on sparse keypoints that provide limited surface constraints. We observe that the true bottleneck lies in the quality of intermediate representations, and that dense pixel-to-surface correspondences can be effectively generated by repurposing pre-trained diffusion models with rich visual priors. We propose DiffProxy, a Stable-Diffusion-based framework trained on large-scale synthetic data with pixel-perfect annotations. A multi-conditional proxy generator predicts dense correspondences from multi-view images, providing uniform surface constraints that enable precise fitting. Hand refinement feeds enlarged hand crops alongside full-body images for fine-grained detail, while test-time scaling exploits diffusion stochasticity to estimate per-pixel uncertainty. Trained only on synthetic data, DiffProxy achieves state-of-the-art results on five diverse real-world benchmarks. Project page: https://wrk226.github.io/DiffProxy.html

人体重建扩散模型多视角稠密对应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。