针对高维小样本表格数据,提出分块子单元扩散生成框架。
BSTabDiff: Block-Subunit Diffusion Priors for High-Dimensional Tabular Data Generation

- 将高维特征分组为少量潜在块,用低维潜变量生成各块。
- 在真实度和稳定性上优于传统表格生成器,尤其在高维小样本场景。
- 适合生物组学等高维稀疏数据的可控合成与隐私保护建模。
高维低样本(HDLSS)表格数据(如组学数据)具有 $n \ll m$ 特征,即样本数远小于特征数。这类数据常呈现强局部相关性、稀疏跨组依赖、重尾非高斯边缘分布、异方差噪声及结构化缺失,导致在 $\mathbb{R}^m$ 中直接建模密度条件恶劣。本文提出 BSTabDiff,一种分块-子单元生成框架:将 $m$ 个观测特征划分为 $M$ 个潜在块($M \ll m$),通过共享的低维子单元变量生成每一块,将全局依赖学习集中在紧凑的块潜空间 $\mathbb{R}^M$,再通过哥德尔依赖机制、灵活的每特征边缘分布和显式缺失建模解码至全特征空间。该框架支持现代深层先验(如扩散模型、归一化流),可在 HDLSS 情况下实现稳定合成与可控基准生成。实验表明,与无结构表格生成器相比,BSTabDiff 在高维合成数据的真实性和稳定性方面表现更优。
原文摘要 · Abstract (English)
High-Dimensional Low-Sample Size (HDLSS) tabular domains (e.g., omics) are characterized by $n \ll m$, where $n$ = number of samples, and $m$ = number of features. Such domains often exhibit strong local correlation groups, sparse cross-group dependencies, heavy-tailed non-Gaussian marginals, heteroscedastic noise, and structured missingness, making direct density learning in $\mathbb{R}^m$ ill-conditioned since $n \ll m$. We propose BSTabDiff, a block-subunit generative framework that partitions the $m$ observed features into $M$ latent blocks ($M \ll m$) and generates each block via a shared low-dimensional subunit variable, concentrating global dependence learning in the compact block-latent space $\mathbb{R}^M$ while decoding to the full feature space with copula-driven dependence, flexible per-feature marginals, and explicit missingness mechanisms. BSTabDiff supports modern deep priors on block latents, including diffusion and normalizing flows, enabling stable synthesis and controllable benchmark generation in the HDLSS regime. Empirically, BSTabDiff produces more realistic and stable high-dimensional synthetic data when compared with unstructured tabular generators on HDLSS data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。