用扩散模型+依赖感知注意力,精准填补微生物组数据空白。
DepMicroDiff: Diffusion-Based Dependency-Aware Multimodal Imputation for Microbiome Data
- 结合扩散模型与依赖感知变压器,捕捉微生物间复杂关系。
- 在多种癌症数据上,相关性达0.712,误差低于现有方法。
- 支持患者元数据条件输入,适合临床微生物研究者使用。
微生物组数据分析对理解宿主健康与疾病至关重要,但其固有的稀疏性和噪声给准确填补带来挑战,影响生物标志物发现等下游任务。现有填补方法,包括最新的扩散模型,常无法捕捉微生物类群间的复杂相互依赖关系,且忽略可辅助填补的上下文元数据。本文提出DepMicroDiff框架,将基于扩散的生成建模与依赖感知变压器(DAT)结合,显式建模微生物间的成对依赖和自回归关系。该模型通过跨多种癌症数据集的变分自编码器(VAE)预训练,并利用大语言模型(LLM)编码患者元数据进行条件化。在TCGA微生物组数据集上的实验表明,DepMicroDiff显著优于当前最优基线,在多种癌症类型中达到最高0.712的皮尔逊相关系数、0.812的余弦相似度,同时均方根误差(RMSE)和平均绝对误差(MAE)更低,证明了其在微生物组填补中的鲁棒性与泛化能力。
原文摘要 · Abstract (English)
Microbiome data analysis is essential for understanding host health and disease, yet its inherent sparsity and noise pose major challenges for accurate imputation, hindering downstream tasks such as biomarker discovery. Existing imputation methods, including recent diffusion-based models, often fail to capture the complex interdependencies between microbial taxa and overlook contextual metadata that can inform imputation. We introduce DepMicroDiff, a novel framework that combines diffusion-based generative modeling with a Dependency-Aware Transformer (DAT) to explicitly capture both mutual pairwise dependencies and autoregressive relationships. DepMicroDiff is further enhanced by VAE-based pretraining across diverse cancer datasets and conditioning on patient metadata encoded via a large language model (LLM). Experiments on TCGA microbiome datasets show that DepMicroDiff substantially outperforms state-of-the-art baselines, achieving higher Pearson correlation (up to 0.712), cosine similarity (up to 0.812), and lower RMSE and MAE across multiple cancer types, demonstrating its robustness and generalizability for microbiome imputation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。