arXiv:2504.16580cs.LGstat.ML2025-04ICML被引 3

用Transformer生成隐变量函数,让扩散模型更高效地表示复杂数据。

Hyper-Transforming Latent Diffusion Models

  • 用Transformer代替MLP做超网络,从隐变量生成INR参数。
  • 在多模态数据上实现更高表达力与可扩展性,无需重训练。
  • 适合想快速适配现有模型到函数表示的研究者。

我们提出一种新型生成框架,将隐式神经表示(INRs)与基于Transformer的超网络整合进隐变量模型中。与依赖受限可扩展性的MLP超网络的以往方法不同,本方法采用Transformer解码器,从隐变量生成INR参数,同时提升表示能力与计算效率。通过用基于Transformer的超网络替换标准解码器,将潜在扩散模型(LDMs)扩展至INR生成任务。该方法可从头训练,也可通过超变换策略实现:仅微调解码器,冻结预训练隐空间,从而无需全量重训即可高效适配现有生成模型至基于INR的表示。我们在多种模态上验证了该方法,结果表明其在可扩展性、表达力和泛化性能上均优于现有基于INR的生成模型。研究建立了一个统一且灵活的结构化函数表示学习框架。

原文摘要 · Abstract (English)

We introduce a novel generative framework for functions by integrating Implicit Neural Representations (INRs) and Transformer-based hypernetworks into latent variable models. Unlike prior approaches that rely on MLP-based hypernetworks with scalability limitations, our method employs a Transformer-based decoder to generate INR parameters from latent variables, addressing both representation capacity and computational efficiency. Our framework extends latent diffusion models (LDMs) to INR generation by replacing standard decoders with a Transformer-based hypernetwork, which can be trained either from scratch or via hyper-transforming: a strategy that fine-tunes only the decoder while freezing the pre-trained latent space. This enables efficient adaptation of existing generative models to INR-based representations without requiring full retraining. We validate our approach across multiple modalities, demonstrating improved scalability, expressiveness, and generalization over existing INR-based generative models. Our findings establish a unified and flexible framework for learning structured function representations.

生成模型隐式表示Transformer扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。