arXiv:2506.20616cs.CV2025-06

将自然物体轮廓转化为创意动物形象,实现视觉想象的自动化。

Shape2Animal: Creative Animal Generation from Natural Silhouettes

  • 用视觉语言模型识别轮廓并联想动物概念
  • 通过扩散模型生成符合形状的动物图像
  • 适合艺术创作与交互媒体设计者使用

人类具有在模糊刺激中感知有意义模式的独特能力,称为拟态现象(pareidolia)。本文提出Shape2Animal框架,模拟这种想象力,将云朵、石头或火焰等自然物体的轮廓重新诠释为合理的动物形态。该自动化框架首先通过开放词汇分割提取物体轮廓,并利用视觉-语言模型推断语义合适的动物概念;随后,借助文本到图像的扩散模型合成符合输入形状的动物图像,并无缝融合到原场景中,生成视觉连贯且空间一致的构图。我们在多样化的现实输入上评估了Shape2Animal,验证了其鲁棒性与创造性潜力。该方法可为视觉叙事、教育内容、数字艺术和交互媒体设计提供新可能。

原文摘要 · Abstract (English)

Humans possess a unique ability to perceive meaningful patterns in ambiguous stimuli, a cognitive phenomenon known as pareidolia. This paper introduces Shape2Animal framework to mimics this imaginative capacity by reinterpreting natural object silhouettes, such as clouds, stones, or flames, as plausible animal forms. Our automated framework first performs open-vocabulary segmentation to extract object silhouette and interprets semantically appropriate animal concepts using vision-language models. It then synthesizes an animal image that conforms to the input shape, leveraging text-to-image diffusion model and seamlessly blends it into the original scene to generate visually coherent and spatially consistent compositions. We evaluated Shape2Animal on a diverse set of real-world inputs, demonstrating its robustness and creative potential. Our Shape2Animal can offer new opportunities for visual storytelling, educational content, digital art, and interactive media design. Our project page is here: https://shape2image.github.io

图像生成创意设计扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。