arXiv:2507.13708cs.CV2025-07被引 2

用多阶段提示优化提升诗歌转图像的信息保留率

PoemTale Diffusion: Minimising Information Loss in Poem to Image Generation with Multi-Stage Prompt Refinement

  • 通过多阶段提示优化增强诗歌文本理解
  • 在P4I数据集上生成更准确传达诗意的图像
  • 无需训练,适合诗歌与艺术创作研究者

近期文本到图像扩散模型在生成真实且多样的视觉内容方面取得显著进展。关键因素在于模型对文本提示的准确解读能力。然而,这些模型在处理创造性表达时仍存在困难,尤其面对复杂、抽象或高度描述性的语言。本文提出PoemTale Diffusion方法,一种无需训练的新型方案,旨在提升诗歌这一特殊创意语言的图像生成效果。诗歌常含多层次、抽象及双关含义,易在转换中丢失信息。该方法通过在语言模型中引入多阶段提示优化循环,增强诗歌文本的可解释性,并改进扩散模型自注意力机制,以生成多张一致图像,共同传达诗作内涵。此外,为推动该领域研究,我们构建了包含1111首诗歌的P4I(PoemForImage)数据集,涵盖线上线下多种来源。邀请诗歌专家进行定性评估。人类与量化评估结果均验证了本方法的有效性,为诗歌到图像生成提供了更强的信息保留新视角。

原文摘要 · Abstract (English)

Recent advancements in text-to-image diffusion models have achieved remarkable success in generating realistic and diverse visual content. A critical factor in this process is the model's ability to accurately interpret textual prompts. However, these models often struggle with creative expressions, particularly those involving complex, abstract, or highly descriptive language. In this work, we introduce a novel training-free approach tailored to improve image generation for a unique form of creative language: poetic verse, which frequently features layered, abstract, and dual meanings. Our proposed PoemTale Diffusion approach aims to minimise the information that is lost during poetic text-to-image conversion by integrating a multi stage prompt refinement loop into Language Models to enhance the interpretability of poetic texts. To support this, we adapt existing state-of-the-art diffusion models by modifying their self-attention mechanisms with a consistent self-attention technique to generate multiple consistent images, which are then collectively used to convey the poem's meaning. Moreover, to encourage research in the field of poetry, we introduce the P4I (PoemForImage) dataset, consisting of 1111 poems sourced from multiple online and offline resources. We engaged a panel of poetry experts for qualitative assessments. The results from both human and quantitative evaluations validate the efficacy of our method and contribute a novel perspective to poem-to-image generation with enhanced information capture in the generated images.

诗歌生成扩散模型提示优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。