arXiv:2410.16840cs.CV2024-10

首个专用于电影海报生成的图文数据集,助力扩散模型产出更精准海报。

MPDS: A Movie Posters Dataset for Image Generation with Diffusion Model

  • 构建37.3万+图文对数据集,含演员图像与标准化描述。
  • 融合自动生成与人工校正,提升提示词与海报内容一致性。
  • 支持个性化海报生成,适合影视、设计领域研究者使用。

电影海报在吸引观众、传递主题和推动市场竞争中至关重要。传统设计耗时费力,而智能生成技术可提升效率与设计质量。然而当前图像生成模型在海报生成上仍表现不佳,主因是缺乏专用海报数据集。本文提出首个面向文本到图像生成的电影海报数据集(MPDS),包含37.3万+图像-文本对及8000+演员图像(覆盖4000+演员)。所有海报描述基于公开电影梗概进行结构化整理,形成“电影梗概提示词”。为进一步增强描述的视觉感知性并缩小与梗概差异,我们利用大规模视觉语言模型自动生成每张海报的视觉感知提示词,再经人工修正并整合至电影梗概提示词。此外,引入“海报标题提示词”以体现海报中的文字元素(如演员名、片名)。针对海报生成,我们设计了多条件扩散框架,输入包括海报提示词、标题提示词和演员图像(用于个性化),通过扩散模型学习实现高质量生成。实验表明,所提MPDS在个性化海报生成中具有显著价值。数据集已开源:https://anonymous.4open.science/r/MPDS-373k-BD3B。

原文摘要 · Abstract (English)

Movie posters are vital for captivating audiences, conveying themes, and driving market competition in the film industry. While traditional designs are laborious, intelligent generation technology offers efficiency gains and design enhancements. Despite exciting progress in image generation, current models often fall short in producing satisfactory poster results. The primary issue lies in the absence of specialized poster datasets for targeted model training. In this work, we propose a Movie Posters DataSet (MPDS), tailored for text-to-image generation models to revolutionize poster production. As dedicated to posters, MPDS stands out as the first image-text pair dataset to our knowledge, composing of 373k+ image-text pairs and 8k+ actor images (covering 4k+ actors). Detailed poster descriptions, such as movie titles, genres, casts, and synopses, are meticulously organized and standardized based on public movie synopsis, also named movie-synopsis prompt. To bolster poster descriptions as well as reduce differences from movie synopsis, further, we leverage a large-scale vision-language model to automatically produce vision-perceptive prompts for each poster, then perform manual rectification and integration with movie-synopsis prompt. In addition, we introduce a prompt of poster captions to exhibit text elements in posters like actor names and movie titles. For movie poster generation, we develop a multi-condition diffusion framework that takes poster prompt, poster caption, and actor image (for personalization) as inputs, yielding excellent results through the learning of a diffusion model. Experiments demonstrate the valuable role of our proposed MPDS dataset in advancing personalized movie poster generation. MPDS is available at https://anonymous.4open.science/r/MPDS-373k-BD3B.

图像生成扩散模型数据集海报设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。