解决人脸生成中对齐、真实感与美感的三难困境
Pareto-Enhanced Portrait Generation: Vision-Aligned Text Supervision for Alignment, Realism, and Aesthetics
- 用视觉对齐文本表示指导图像生成,不增加推理开销
- 在保持模型泛化能力下,同时提升对齐、真实感和美感
- 适合追求高质量人脸生成的开发者与研究者
文本到图像扩散模型在生成人像时面临严重三难困境:文本-图像对齐、照片真实感与人类感知美感彼此制约。监督微调(SFT)虽能提升真实感,但常导致过拟合、破坏预训练图像先验,并损害对齐或美感。为此,我们提出一种针对多模态扩散变换器(MM-DiT)的特征监督范式。具体而言,引入轻量级跨模态对齐机制,从SigLIP 2隐式提取多粒度视觉对齐文本表征,并在训练阶段对MM-DiT的图像分支施加监督,实现零额外推理开销。该方法注入视觉对齐文本引导,同时保留基座模型原始泛化能力,避免SFT带来的退化。此外,直接从预训练视觉基础模型中挖掘隐式多粒度审美信号,以优化人类感知美感。大量实验表明,本方法显著推动了帕累托前沿,在文本-图像对齐、真实感与美学感知上实现协同提升。
原文摘要 · Abstract (English)
Text-to-image diffusion models often face a severe trilemma in human portrait generation: text-image alignment, photorealism, and human-perceived aesthetics inherently inhibit one another. Supervised Fine-Tuning (SFT) is an effective method for enhancing the photorealism of image generation. However, it often leads to overfitting to the training dataset, corrupts pre-trained image priors, and degrades alignment or aesthetics. To break this bottleneck, we propose a feature supervision paradigm for Multimodal Diffusion Transformers (MM-DiT). Specifically, we introduce a lightweight cross-modal alignment mechanism that implicitly extracts multi-granularity vision-aligned text representations from SigLIP 2 and applies supervision to the image branch of MM-DiT during the training stage, with zero extra inference overhead. Our method injects vision-aligned text guidance while preserving the base model's original generalization, avoiding degradation caused by SFT. Furthermore, our method directly mines implicit multi-granularity aesthetic signals from pre-trained vision foundation models to optimize human-perceived aesthetics. Extensive experiments on MM-DiTs show that our method pushes the Pareto frontier and achieves synergistic improvements across text-image alignment, photorealism, and human-perceived aesthetics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。