arXiv:2410.13034q-bio.NCcs.CV2024-10被引 3

用Stable Diffusion生成连续变化的高分辨率自然图像,用于感知研究。

Synthesis and Perceptual Scaling of High Resolution Naturalistic Images Using Stable Diffusion

  • 基于Stable Diffusion XL生成类别内渐变的逼真图像。
  • 构建108个物体场景,每类10个变体,按感知相似性排序。
  • 适用于视觉感知、注意与记忆中的渐进式刺激研究。

自然场景对视觉感知研究至关重要,但控制其感知与语义特性颇具挑战。以往研究多聚焦于物理差异显著的离散图像集,而实际中常需沿连续维度评估自然图像表征。传统方法通过图像插值生成连续变化,但主要依赖低层物理特征,易产生语义模糊结果。近期研究采用生成对抗网络(GAN)在类别内生成连续感知变化。本文改用文本到图像扩散模型Stable Diffusion XL,生成可自由定制的摄影级逼真图像集,实现类别内渐进过渡,每个图像为提示类别下的独特实例。我们以6类物体场景为例,每类生成10个变体,共108张图像。先通过感知相似性模型LPIPS估计排序,再经大规模在线人类样本验证。后续实验表明该排序能预测工作记忆中的混淆程度。本图像集适用于研究自然刺激在视觉感知、注意与记忆中的渐进编码机制。

原文摘要 · Abstract (English)

Naturalistic scenes are of key interest for visual perception, but controlling their perceptual and semantic properties is challenging. Previous work on naturalistic scenes has frequently focused on collections of discrete images with considerable physical differences between stimuli. However, it is often desirable to assess representations of naturalistic images that vary along a continuum. Traditionally, perceptually continuous variations of naturalistic stimuli have been obtained by morphing a source image into a target image. This produces transitions driven mainly by low-level physical features and can result in semantically ambiguous outcomes. More recently, generative adversarial networks (GANs) have been used to generate continuous perceptual variations within a stimulus category. Here we extend and generalize this approach using a different machine learning approach, a text-to-image diffusion model (Stable Diffusion XL), to generate a freely customizable stimulus set of photorealistic images that are characterized by gradual transitions, with each image representing a unique exemplar within a prompted category. We demonstrate the approach by generating a set of 108 object scenes from 6 categories. For each object scene, we generate 10 variants that are ordered along a perceptual continuum. This ordering was first estimated using a machine learning model of perceptual similarity (LPIPS) and then subsequently validated with a large online sample of human participants. In a subsequent experiment we show that this ordering is also predictive of confusability of stimuli in a working memory experiment. Our image set is suited for studies investigating the graded encoding of naturalistic stimuli in visual perception, attention, and memory.

图像生成扩散模型感知研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。