用对齐张量空间实现高效高分辨率生成,无需预训练压缩模型。
DiffATS: Diffusion in Aligned Tensor Space

- 基于张量分解与正交对齐,构建数据自适应的紧凑张量基元。
- 在图像、视频和偏微分方程解上实现3.9至210倍压缩,生成效果强。
- 理论证明表示空间拓扑保真,适合无预训练压缩的生成任务。
高维时空场的直接扩散建模计算成本高。参数高效的基元通过紧凑参数集表示高维数据。本文提出一种无需预训练压缩自编码器的数据依赖张量基元构造方法。从Tucker分解出发,利用核心张量和模式因子捕捉低秩多线性结构。但Tucker因子不唯一:同一张量可由不同旋转因子表示,导致生成建模困难。为此,我们采用正交Procrustes(OP)对齐策略,从数据中选取中位锚矩阵,并对因子矩阵进行对齐以消除规范歧义。由此得到可直接解码的矩阵和张量Grassmann基元,具有紧凑性、数据自适应性和显式多线性重构能力。理论上证明所提基元映射为低秩张量与其基元空间间的同胚,确保表示非退化且拓扑忠实。在此基础上,提出扩散在对齐张量空间(DiffATS)的生成框架,直接在对齐基元上训练扩散模型。在图像、视频和偏微分方程解上,DiffATS实现了强的无条件与条件生成性能,同时将原始数据压缩3.9×至210×,且无需任何预训练深度压缩自编码器。
原文摘要 · Abstract (English)
Direct diffusion modeling of high-resolution spatiotemporal fields is computationally challenging. Parameter-efficient primitives address this by representing high-dimensional data with a compact set of parameters. In this paper, we construct data-dependent tensor primitives without pretrained compression autoencoders. Our construction starts from Tucker decomposition, which captures low-rank multilinear structure through a core tensor and mode-wise factors. However, Tucker factors are non-unique: the same tensor can be represented by different rotated factors, which complicates generative modeling. We address this issue with orthogonal Procrustes (OP) alignment. Specifically, we select medoid anchor matrices from the data and align the factor matrices to resolve the gauge ambiguity. This yields matrix Grassmannian primitives and tensor Grassmannian primitives that are compact, data-adaptive, and directly decodable by explicit multilinear reconstruction. Theoretically, we prove that the proposed primitive maps are homeomorphisms between low-rank tensors and their corresponding primitive spaces, certifying that the representations are non-degenerate and topologically faithful. Building on these primitives, we propose *Diffusion in Aligned Tensor Space* (DiffATS), a generative framework that trains diffusion models directly on aligned tensor primitives. Across images, videos, and PDE solutions, DiffATS achieves strong unconditional and conditional generation performance while compressing original data by $3.9\times$ to $210\times$, without relying on any pretrained deep compression autoencoders.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。