用3D关键点和相机参数控制真人图像生成,避免2D投影失真。
PoseCraft: Tokenized 3D Body Landmark and Camera Conditioning for Photorealistic Human Image Synthesis
- 将3D人体关键点和相机参数转为离散条件令牌,注入扩散模型。
- 在大幅姿态和视角变化下仍保持3D语义一致性,细节还原更真实。
- 适合需要精确控制姿态与视角的虚拟人、VR和影视制作场景。
数字化人类并生成可精确控制3D姿态与相机视角的逼真人像,是虚拟现实、远程呈现和娱乐的核心需求。现有基于骨骼绑定的方法需大量手动绑定或模板适配,神经体素方法则依赖标准模板且需为每个新姿态重新优化。本文提出PoseCraft,一种基于3D关键点与相机外参离散令牌的扩散框架:不依赖2D渲染图像作为控制输入,而是将稀疏3D关键点和相机外参编码为离散条件令牌,通过交叉注意力注入扩散过程。该方法避免了大姿态与视角变化下的2D重投影模糊问题,有效保留3D语义,生成的图像能忠实还原身份与外观。为支持大规模训练与评估,我们还构建了GenHumanRF数据生成流程,从体素重建中合成多样化监督信号。实验表明,PoseCraft在感知质量上显著优于主流扩散方法,在最新体素渲染方法的指标上达到相当或更好表现,尤其在衣物与头发细节保留上更具优势。
原文摘要 · Abstract (English)
Digitizing humans and synthesizing photorealistic avatars with explicit 3D pose and camera controls are central to VR, telepresence, and entertainment. Existing skinning-based workflows require laborious manual rigging or template-based fittings, while neural volumetric methods rely on canonical templates and re-optimization for each unseen pose. We present PoseCraft, a diffusion framework built around tokenized 3D interface: instead of relying only on rasterized geometry as 2D control images, we encode sparse 3D landmarks and camera extrinsics as discrete conditioning tokens and inject them into diffusion via cross-attention. Our approach preserves 3D semantics by avoiding 2D re-projection ambiguity under large pose and viewpoint changes, and produces photorealistic imagery that faithfully captures identity and appearance. To train and evaluate at scale, we also implement GenHumanRF, a data generation workflow that renders diverse supervision from volumetric reconstructions. Our experiments show that PoseCraft achieves significant perceptual quality improvement over diffusion-centric methods, and attains better or comparable metrics to latest volumetric rendering SOTA while better preserving fabric and hair details.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。