arXiv:2409.15512cs.CVcs.AI2024-09

统一表征让文本与像素图像生成更连贯。

PixelBytes: Catching Unified Embedding for Multimodal Generation

  • 用PxBy嵌入技术融合多模态数据,支持双向序列建模。
  • 在专有宝可梦数据集上实现连贯的图文序列生成。
  • 适合研究统一多模态生成的学者和工程师。

本文提出PixelBytes嵌入,一种统一的多模态表征学习方法。该方法将多样化输入整合为单一连贯表示,促进文本与像素化图像的涌现式序列生成。受Image Transformers、PixelCNN和Mamba-Bytes等先进序列模型启发,我们探索了循环神经网络(RNN)、状态空间模型(SSMs)及基于注意力的模型,重点研究双向处理与创新的PxBy嵌入技术。在专用PixelBytes Pok{é}mon数据集上的实验表明,结合PxBy嵌入与卷积层的双向序列模型能够生成连贯的多模态序列。本工作推动了能统一理解与生成多模态数据的集成式AI模型的发展。

原文摘要 · Abstract (English)

This report introduces PixelBytes Embedding, a novel approach for unified multimodal representation learning. Our method captures diverse inputs in a single, cohesive representation, enabling emergent properties for multimodal sequence generation, particularly for text and pixelated images. Inspired by state-of-the-art sequence models such as Image Transformers, PixelCNN, and Mamba-Bytes, PixelBytes aims to address the challenges of integrating different data types. We explore various model architectures, including Recurrent Neural Networks (RNNs), State Space Models (SSMs), and Attention-based models, focusing on bidirectional processing and our innovative PxBy embedding technique. Our experiments, conducted on a specialized PixelBytes Pok{é}mon dataset, demonstrate that bidirectional sequence models with PxBy embedding and convolutional layers can generate coherent multimodal sequences. This work contributes to the advancement of integrated AI models capable of understanding and generating multimodal data in a unified manner.

多模态生成统一表征序列建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。