arXiv:2606.17431cs.CV2026-06被引 1

让机器从自然轮廓中生成动物艺术,模仿人类的联想力

Visual Retrieval-Augmented Generation for Silhouette-Guided Animal Art

论文配图:Visual Retrieval-Augmented Generation for Silhouette-Guided Animal Art
图 1 · 摘自论文原文
  • 用2.8万张动物轮廓检索相似形状,作为生成参考
  • 通过控制网与IP适配器实现轮廓引导的扩散生成
  • 适合对创意生成和视觉联想感兴趣的设计师

生成式AI已能绘制逼真或艺术化图像,但在解读模糊形状方面仍受限于人类创造力。这一现象源于帕莱多利亚——人类能从云朵、石头或树叶等随机图案中感知有意义形态。为计算复现此想象过程,我们提出视觉检索增强生成(Visual-RAG)框架,直接从自然轮廓生成动物艺术。方法从包含28,586张高质量轮廓的精选语料库中检索结构相似的动物形状,并将其作为参考示例,结合ControlNet与IP-Adapter引导扩散模型生成。消融实验表明,使用RANSAC的形状上下文提供最准确对齐,而移除形状标准化后内点率仅降至13.4%,凸显结构保真度的重要性。12名参与者进行用户研究,评估美学、轮廓保真度与整体印象。结果表明,Visual-RAG虽能提供合理解释,但感知冲击力仍有提升空间。本工作为计算帕莱多利亚奠定基础,展示机器如何参与早期创意发现阶段。

原文摘要 · Abstract (English)

Generative AI has advanced the ability to render photorealistic or artistic images, yet it remains limited in a key aspect of human creativity: interpreting ambiguous shapes. This phenomenon, rooted in pareidolia, allows humans to perceive meaningful forms in random patterns such as clouds, stones, or leaves. To computationally replicate this imaginative process, we introduce Visual Retrieval-Augmented Generation (Visual-RAG), a framework that generates animal art directly from natural silhouettes. Our method retrieves structurally similar animal shapes from a curated corpus of 28,586 high-quality silhouettes and uses them as reference exemplars to guide diffusion-based generation with ControlNet and IP-Adapter. Ablation studies confirm that shape Context with RANSAC provides the most accurate alignment, while removing shape standardization reduces the inlier ratio to just 13.4\%, underscoring the importance of structural fidelity in Visual-RAG. A user study with 12 participants evaluated the outputs in terms of aesthetics, silhouette fidelity, and overall impression. Results reveal that while Visual-RAG provides plausible interpretations, challenges remain in achieving high perceptual impact. This work lays the foundation for computational pareidolia, showing how machines can contribute to the early stages of imaginative discovery.

图像生成轮廓引导创意设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。