arXiv:2410.17970cs.NEcs.LG2024-10被引 68

用光学实现快速低功耗图像生成,无需计算即可出新图。

Optical Generative Models

  • 先用浅层数字编码器生成相位种子,再由全光学解码器生成图像。
  • 在MNIST、Fashion MNIST等数据集上生成效果媲美数字模型。
  • 适合追求超低功耗与极速生成的AI内容应用。

生成模型广泛应用于图像、视频、音乐合成、自然语言处理及分子设计等领域。随着数字生成模型规模增大,如何实现快速且节能的可扩展推理成为挑战。本文提出受扩散模型启发的光学生成模型:通过一个浅层快速数字编码器将随机噪声映射为相位图案,作为目标数据分布的光学生成种子;随后,联合训练的基于自由空间的可重构光学解码器全光处理这些种子,生成符合目标数据分布的新图像(从未见过)。除照明功率和浅层编码器生成随机种子外,图像生成过程不消耗计算功耗。实验中实现了基于手写数字、服装产品、蝴蝶和人脸数据集(MNIST、Fashion MNIST、Butterflies-100、Celeb-A)的单色及多色新图像生成,整体性能与数字神经网络生成模型相当。通过可见光在单帧内完成手写数字与服装产品的光学生成,验证了可行性。该方法为实现能效高、可扩展、快速的推理任务开辟新路径,进一步挖掘光学与光子学在人工智能生成内容中的潜力。

原文摘要 · Abstract (English)

Generative models cover various application areas, including image, video and music synthesis, natural language processing, and molecular design, among many others. As digital generative models become larger, scalable inference in a fast and energy-efficient manner becomes a challenge. Here, we present optical generative models inspired by diffusion models, where a shallow and fast digital encoder first maps random noise into phase patterns that serve as optical generative seeds for a desired data distribution; a jointly-trained free-space-based reconfigurable decoder all-optically processes these generative seeds to create novel images (never seen before) following the target data distribution. Except for the illumination power and the random seed generation through a shallow encoder, these optical generative models do not consume computing power during the synthesis of novel images. We report the optical generation of monochrome and multi-color novel images of handwritten digits, fashion products, butterflies, and human faces, following the data distributions of MNIST, Fashion MNIST, Butterflies-100, and Celeb-A datasets, respectively, achieving an overall performance comparable to digital neural network-based generative models. To experimentally demonstrate optical generative models, we used visible light to generate, in a snapshot, novel images of handwritten digits and fashion products. These optical generative models might pave the way for energy-efficient, scalable and rapid inference tasks, further exploiting the potentials of optics and photonics for artificial intelligence-generated content.

光学生成低功耗全光计算图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。