让大模型生成可控3D人像,自由调节外貌与姿态。
Disentangling 3D from Large Vision-Language Models for Controlled Portrait Generation
- 用2D参考帧解耦3D特征,结合可变形神经三平面表示。
- 引入雅可比正则化,有效降低大模型嵌入空间噪声影响。
- 仅需2D数据即可实现3D姿态与文本属性的独立控制,适合内容创作者使用。
本文研究如何从大型视觉语言模型中解耦3D信息,以实现生成可控的3D人像。该方法支持自由文本控制年龄、发型、眼镜等外观属性,以及面部表情和相机位姿的3D几何控制。在不依赖额外配对标签的前提下,基于预训练的大型视觉语言模型(如CLIP)和预定义的3D可变形模型(FLAME),通过将可变形神经三平面表示归一化到2D参考帧来实现解耦。然而,大模型嵌入空间中的显著噪声会引入无关特征,损害生成质量和多样性。为此,本文提出一种高效的随机近似计算的雅可比正则化方法加以克服。相比现有方法,本方案可在保持一致性的同时,实现文本与3D控制的独立调节。该方法使创作者仅用2D人脸数据即可控制3D生成器,无需大规模标注或训练大模型。
原文摘要 · Abstract (English)
We consider the problem of disentangling 3D from large vision-language models, which we show on generative 3D portraits. This allows free-form text control of appearance attributes like age, hair style, and glasses, and 3D geometry control of face expression and camera pose. In this setting, we assume we use a pre-trained large vision-language model (LVLM; CLIP) to generate from a smaller 2D dataset with no additional paired labels and with a pre-defined 3D morphable model (FLAME). First, we disentangle using canonicalization to a 2D reference frame from a deformable neural 3D triplane representation. But another form of entanglement arises from the significant noise in the LVLM's embedding space that describes irrelevant features. This damages output quality and diversity, but we overcome this with a Jacobian regularization that can be computed efficiently with a stochastic approximator. Compared to existing methods, our approach produces portraits with added text and 3D control, where portraits remain consistent when either control is changed. Broadly, this approach lets creators control 3D generators on their own 2D face data without needing resources to label large data or train large models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。