将梦境文字描述转化为连贯的四格图像序列。
Coherence-Oriented Dream Scene Visualisation

- 用大语言模型拆分梦境为四个时间顺序片段。
- 生成四幅图像并保持视觉连贯性,错误图像可重生成。
- 适合心理学研究与梦境可视化,效果经多模型验证。
梦境情感强烈但难以表达。本文提出梦境场景可视化系统(DSV),将文字描述的梦境转换为由四幅图像组成的时序序列。首先利用大语言模型将梦境描述划分为四个时间顺序部分;随后,文本到图像模型为每部分生成图像,并确保序列间视觉连贯;若某图与文本不符,DSV会进行重生成。我们在DreamBank数据集上对50个梦境进行了评估,通过CLIP、DINOv2和Qwen2-VL等视觉-语言模型进行客观度量,报告了质量、保真度与连贯性结果。
原文摘要 · Abstract (English)
Dreams can be emotionally intense but difficult to communicate. We describe the Dream Scene Visualiser (DSV) system which turns written dream descriptions into a temporal sequence of four panel images visualising the dream. This starts with a large language model prompted to split a dream description into four chronological parts. Then a text-to-image model produces images for each part with visual coherence maintained across the sequence, and DSV regenerates any image not suitably matching the text. We evaluate DSV over 50 visualisations from dream descriptions in DreamBank, and report quality, fidelity and coherence results via objective measures employing the CLIP, DINOv2 and Qwen2-VL vision-language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。