提出新方法精准发现纵向组学数据中的隐藏模式,助力疾病分型研究。
Sparse Functional Singular Value Decomposition for Biclustering and Triclustering Longitudinal Data

- 基于稀疏函数SVD框架,同步筛选个体、特征与时间片段
- 在高维稀疏数据上表现优于现有方法,识别出三类肠道菌群关联集群
- 适用于复杂疾病异质性分析,适合生物医学研究人员使用
识别复杂疾病如炎症性肠病(IBD)的亚型,需挖掘纵向组学数据中的潜在规律。然而,此类数据通常维度高、采样稀疏且时间点不规则,传统(双)聚类和函数数据分析方法难以应对。本文提出Tri-SfSVD,一种统一的稀疏函数奇异值分解框架,用于发现纵向数据中的双聚类与三聚类结构。不同于依赖人工填补或强形状假设的现有方法,Tri-SfSVD将连续轨迹估计与个体、变量及时间子区间的同时选择整合于单一优化框架中。通过在个体、变量和时间子区域上施加稀疏惩罚,该方法直接处理观测数据,揭示个体、个体-特征以及个体-特征-时间层面的局部结构。大量模拟实验表明,其在高维情形下优于现有方法。应用于IBD多组学数据,识别出三个双聚类:分别关联样本聚类与特定临床特征的微生物通路群,涉及特定细菌分类群;应用于多通道脑电图(EEG)数据,识别出三个三聚类:关联样本聚类与酒精相关表型的局域脑活动模式,包括同一空间区域内的时序子区间差异。
原文摘要 · Abstract (English)
Identifying subtypes of complex conditions, such as Inflammatory Bowel Disease (IBD), often requires capturing latent patterns in longitudinal omics data. However, these data are typically high-dimensional, sparsely sampled, and irregularly observed over time, posing substantial challenges for conventional (bi)clustering and functional data analysis methods. We propose Tri-SfSVD, a unified sparse functional Singular Value Decomposition framework for discovering biclusters and triclusters in longitudinal data. Unlike existing functional biclustering methods that rely on ad hoc imputation or enforce restrictive shape-homogeneity assumptions, Tri-SfSVD integrates continuous trajectory estimation with simultaneous subject, feature, and temporal selection within a single optimization framework. By imposing sparse penalties across subjects, variables, and temporal subregions, the proposed method works directly on observed data to uncover localized structures at the subject, subject-feature, and subject-feature-time levels. Extensive simulations demonstrate that Tri-SfSVD outperforms existing approaches in high-dimensional settings. Applied to IBD multi-omics data, the method identified three biclusters linking sample clusters with distinct IBD-related clinical characteristics to microbial pathway groups associated with specific bacterial taxa, providing interpretable subject-pathway associations for characterizing disease heterogeneity. Applied to multi-channel EEG data, the method identified three triclusters linking sample clusters with distinct alcohol-related phenotypes to localized brain activity patterns, including subgroup differences separated by temporal subregions within the same spatial region.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。