用提示调优生成诗意图像,让诗的意境变成画面。
Poetry in Pixels: Prompt Tuning for Poem Image Generation via Diffusion Models
- 通过提示调优让扩散模型理解诗歌深层含义。
- 提出PoeKey算法提取情绪、意象、主题三要素生成指令。
- 构建1001首童诗图文数据集,提升生成多样性。
文本到图像生成在文学作品,尤其是诗歌应用中面临显著挑战,因诗歌意义常超越字面。为此,我们提出PoemToPixel框架,旨在生成体现诗歌内在意义的视觉图像。该方法在图像生成框架中引入提示调优,确保生成图像与诗意内容高度契合。此外,提出PoeKey算法,从诗歌中提取情绪、视觉元素和主题三要素,形成指令输入扩散模型以生成对应图像。为拓展诗歌数据集在不同体裁和年龄层的多样性,我们构建了包含1001首儿童诗及其配图的MiniPo多模态数据集。结合PoemSum数据集,我们对基于PoemToPixel框架的图像生成进行了定量与定性评估。结果表明该方法有效,为从文学文本生成图像提供了新视角。
原文摘要 · Abstract (English)
The task of text-to-image generation has encountered significant challenges when applied to literary works, especially poetry. Poems are a distinct form of literature, with meanings that frequently transcend beyond the literal words. To address this shortcoming, we propose a PoemToPixel framework designed to generate images that visually represent the inherent meanings of poems. Our approach incorporates the concept of prompt tuning in our image generation framework to ensure that the resulting images closely align with the poetic content. In addition, we propose the PoeKey algorithm, which extracts three key elements in the form of emotions, visual elements, and themes from poems to form instructions which are subsequently provided to a diffusion model for generating corresponding images. Furthermore, to expand the diversity of the poetry dataset across different genres and ages, we introduce MiniPo, a novel multimodal dataset comprising 1001 children's poems and images. Leveraging this dataset alongside PoemSum, we conducted both quantitative and qualitative evaluations of image generation using our PoemToPixel framework. This paper demonstrates the effectiveness of our approach and offers a fresh perspective on generating images from literary sources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。