arXiv:2608.23410cs.CVcs.LG2026-08

用新型自回归架构实现高分辨率人脸多视角逼真生成

Photorealistic Novel View Synthesis of Human Faces using Next-Scale Transformers

论文配图:Photorealistic Novel View Synthesis of Human Faces using Next-Scale Transformers
图 1 · 摘自论文原文
  • 基于下一尺度自回归,单次前向输出多视角高清图像
  • 在真实人脸数据集上训练,生成结果更保真且视图间一致性更强
  • 无需2D预训练,小数据也能高效收敛,适合真实场景应用

高分辨率下的人脸多视角逼真合成仍具挑战,需同时保持身份一致、细节清晰与几何连贯。本文基于下一尺度自回归框架,通过支持更高图像分辨率、多视角输出及单次前向传播中的强跨视角一致性,实现面向人像的视图合成。模型在涵盖多样身份与服饰的合成人脸数据集上训练。与扩散模型不同,该方法无需2D预训练,得益于其下一尺度结构,可利用低分辨率通用预训练,仅在最后阶段使用全尺寸特定任务图像。这使模型在较少特定数据下即可收敛,从而使用更小但更真实的训练数据集。实验表明,生成图像更清晰逼真,多视角同步生成提升了视图间一致性。此外,本方法与现有基于Transformer的像素对齐3D高斯提升模型结合,可生成精确且逼真的3D人脸模型。

原文摘要 · Abstract (English)

Photorealistic novel view synthesis of people remains challenging at high spatial resolutions and across multiple target cameras, where preserving identity, fine appearance details, and geometric coherence is critical. We build on the next-scale autoregressive paradigm and adapt it for human-centric view synthesis by enabling higher image resolutions, multi-view outputs and stronger cross-view consistency in a single forward pass. We train on a synthetic dataset of human faces spanning diverse identities and apparel. Contrary to diffusion models, this paradigm does not need 2D pre-training and, thanks to its next-scale architecture, it benefits from lower-resolution, general-purpose pre-trainings, with the full-sized purpose-specific images being used only in the last training stages. This enables our architecture to converge with a smaller amount of purpose-specific training data, allowing us to use a smaller but more realistic training dataset. The resulting model produces sharp and realistic views, with the option to synthesize multiple novel viewpoints simultaneously for improved agreement across views. Empirically, we observe gains in perceptual fidelity and cross-view coherence on human subjects, demonstrating that next-scale autoregression is an effective backbone for scalable, multi-output human view synthesis. We also couple our pipeline with an existing transformer-based model for pixel-aligned 3D gaussian lifting from multi-view facial inputs, resulting in accurate and photorealistic 3D models of human faces.

人脸生成自回归模型多视角合成3D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。