分离成像因素,让文本生成图像更精准可控。
MULTI: Disentangling Camera Lens, Sensor, View, and Domain for Novel Image Generation

- 通过两阶段训练分离镜头、传感器、视角等成像因素。
- 在新基准DF-RICO上生成图像质量显著提升。
- 适合需要精细控制图像风格的视觉生成研究者。
近期文本到图像模型虽能生成高质量图像,但文本描述模糊导致在需要特定风格或物体时控制力不足。现有工作多关注图像内容,忽视了镜头、传感器、视角及场景领域等成像因素。本文提出成像因素解耦新挑战,并引入MULTI方法:第一阶段学习通用因素,第二阶段提取数据集特异性因素。该设计支持扩展现有数据集和新因素组合,缩小分布差异,同时支持针对特定因素修改与基于ControlNets的图像到图像生成。在新提出的DF-RICO基准上的评估验证了MULTI的有效性,凸显了成像因素解耦作为新兴研究方向的重要性。
原文摘要 · Abstract (English)
Recent text-to-image models produce high-quality images, yet text ambiguity hinders precise control when specific styles or objects are required. There have been a number of recent works dealing with learning and composing multiple objects and patterns. However, current work focuses almost entirely on image content, overlooking imaging factors such as camera lens, sensor types, imaging viewpoints, and scenes' domain characteristics. We introduce this new challenge as Imaging Factor Disentanglement and show limitations of current approaches in the regime. We, therefore, propose the new method Multi-factor disentanglement through Textual Inversion (MULTI). It consists of two stages: in the first stage, we learn general factors, and in the second stage, we extract dataset-specific ones. This setup enables the extension of existing datasets and novel factor combinations, thereby reducing distribution gaps. It further supports modifications of specific factors and image-to-image generation via ControlNets. The evaluation on our new DF-RICO benchmark demonstrates the effectiveness of MULTI and highlights the importance of Factor Disentanglement as a new direction of research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。