用视觉语言模型提升图像生成中成像因子的解耦能力
X-MULTI: VLM-based Imaging Factor Disentanglement for Factor-Aware Image Synthesis

- 借助预训练VLM监督未见的成像因子组合,实现更优解耦
- 新方法在未见组合上因子对齐率显著优于MULTI
- 提出新评估指标I-FAA,有效减少评估偏差,适合研究生成可控性
文本到图像生成中的成像因子解耦旨在独立控制相机镜头类型、传感器类型、视角和场景域等图像采集属性,以实现组合泛化。这使模型能合成训练数据中未出现的因子组合,如将鱼眼镜头与事件传感器配对。现有方法MULTI通过可学习的因子特异性嵌入实现解耦,并引入因子对齐准确率(FAA)评估解耦质量。本文发现两大局限:其一,MULTI的像素级重建目标仅监督已观测的因子组合,缺乏对新组合的直接训练信号;为此提出X-MULTI,利用预训练视觉语言模型(VLM)在训练中监督新组合生成。其二,FAA指标存在严重跨因子相关性泄漏,误导真实解耦质量评估;为此提出改进版I-FAA,采用因子特异性增强策略打破相关性,实现更严格评估。实验表明,X-MULTI在未见组合上的因子对齐表现优于MULTI;同时,I-FAA显著降低相关性泄漏,提供更可靠的评估结果。
原文摘要 · Abstract (English)
Imaging factor disentanglement in text-to-image generation aims to independently control image acquisition properties such as types of camera lenses, sensor types, viewpoints, and domains to enable combinatorial generalization. This should let the model synthesize novel factor combinations unobserved in the training data, such as pairing a fisheye lens with an event sensor never observed in training data. Recent work, MULTI, introduced learnable, factor-specific embeddings to disentangle imaging factors, along with the Factor Alignment Accuracy (FAA) metric to evaluate disentanglement quality. We identify and address two independent limitations. First, MULTI's pixel-level reconstruction objective supervises the model only on observed imaging factor combinations, providing no direct training signal for novel combinations. We therefore propose X-MULTI, which uses a pretrained vision-language model (VLM) to supervise novel factor combinations synthesized during training. Second, we show the FAA metric exhibits severe cross-factor correlation leakage, misrepresenting true disentanglement quality. We therefore propose Improved-FAA (I-FAA), which employs factor-specific augmentation strategies to break these correlations and enables more rigorous evaluation. Experiments demonstrate that X-MULTI achieves improved factor alignment on novel combinations compared to MULTI. Moreover, we show that correlation leakage in FAA distorts the evaluation of true factor disentanglement and I-FAA reduces this leakage and therefore provides a more robust assessment of factor alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。