用双条件扩散模型提升图像美感,让编辑更懂审美。
Enhancing Image Aesthetics with Dual-Conditioned Diffusion Models Guided by Multimodal Perception
- 引入多模态感知,将模糊审美指令转为明确指导。
- 在弱配对数据上训练,显著提升美学评分与内容一致性。
- 适合需要智能美化图像的设计师和AI应用开发者。
图像美学增强旨在识别图像的美学缺陷并进行相应编辑,极具挑战性,需模型具备创造力与审美感知能力。尽管近期图像编辑模型在可控性和灵活性上取得进展,但在美学增强方面仍表现不佳。主要挑战有二:一是按审美指令执行编辑困难;二是缺乏内容一致但美学差异明显的“完美配对”图像。本文提出基于扩散模型的双监督图像美学增强方法(DIAE),首先引入多模态美学感知(MAP),通过标准化的多维度美学指令和文本-图像对生成的一致性控制信号,将模糊指令转化为明确引导;其次,为缓解数据稀缺问题,构建了包含相同语义但不同美学质量的“不完美配对”数据集IIAEData,并设计双分支监督框架以利用其弱匹配特性。实验表明,DIAE优于基线模型,在美学评分和内容一致性上均取得显著提升。
原文摘要 · Abstract (English)
Image aesthetic enhancement aims to perceive aesthetic deficiencies in images and perform corresponding editing operations, which is highly challenging and requires the model to possess creativity and aesthetic perception capabilities. Although recent advancements in image editing models have significantly enhanced their controllability and flexibility, they struggle with enhancing image aesthetic. The primary challenges are twofold: first, following editing instructions with aesthetic perception is difficult, and second, there is a scarcity of "perfectly-paired" images that have consistent content but distinct aesthetic qualities. In this paper, we propose Dual-supervised Image Aesthetic Enhancement (DIAE), a diffusion-based generative model with multimodal aesthetic perception. First, DIAE incorporates Multimodal Aesthetic Perception (MAP) to convert the ambiguous aesthetic instruction into explicit guidance by (i) employing detailed, standardized aesthetic instructions across multiple aesthetic attributes, and (ii) utilizing multimodal control signals derived from text-image pairs that maintain consistency within the same aesthetic attribute. Second, to mitigate the lack of "perfectly-paired" images, we collect "imperfectly-paired" dataset called IIAEData, consisting of images with varying aesthetic qualities while sharing identical semantics. To better leverage the weak matching characteristics of IIAEData during training, a dual-branch supervision framework is also introduced for weakly supervised image aesthetic enhancement. Experimental results demonstrate that DIAE outperforms the baselines and obtains superior image aesthetic scores and image content consistency scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。