arXiv:2410.07890stat.MLcs.LG2024-10被引 1

用新方法找出神经疾病中不同患者亚群的隐藏致病因素

Identifying latent disease factors differently expressed in patient subgroups using group factor analysis

  • 通过带正则化霍舍尔先验的稀疏组因子分析,识别多模态数据中的潜在关联
  • 在GENFI数据集上发现两类基因亚群特异的潜在致病因子,且与脑结构和行为评分相关
  • 适合研究神经精神疾病异质性、需整合多源数据的科研人员使用

本研究提出一种新方法——带正则化霍舍尔先验的稀疏组因子分析(sparse GFA),用于揭示神经与精神疾病中不同患者亚群的共性与特异性潜变量。该方法基于概率编程实现,可识别多模态数据中在样本亚群中差异表达的潜在关联。合成数据实验表明其能准确推断潜变量与模型参数。在遗传性额颞叶痴呆倡议(GENFI)数据集上的应用中,该方法成功识别出在不同遗传亚群间差异表达的潜变量,区分了亚群特异性和共性潜变量。这些潜变量捕捉了脑结构与非影像学变量(如行为评估量表和疾病严重程度)之间的关联,揭示了不同基因亚群的疾病特征。特别地,两个潜变量在更同质的前颗粒蛋白(GRN)和微管相关蛋白tau(MAPT)突变亚群中表现更显著,验证了方法识别亚群特异性特征的能力。结果表明,sparse GFA具有整合多模态数据并发现可解释潜变量的潜力,有助于改善神经精神疾病的患者分型与表征。

原文摘要 · Abstract (English)

In this study, we propose a novel approach to uncover subgroup-specific and subgroup-common latent factors addressing the challenges posed by the heterogeneity of neurological and mental disorders, which hinder disease understanding, treatment development, and outcome prediction. The proposed approach, sparse Group Factor Analysis (GFA) with regularised horseshoe priors, was implemented with probabilistic programming and can uncover associations (or latent factors) among multiple data modalities differentially expressed in sample subgroups. Synthetic data experiments showed the robustness of our sparse GFA by correctly inferring latent factors and model parameters. When applied to the Genetic Frontotemporal Dementia Initiative (GENFI) dataset, which comprises patients with frontotemporal dementia (FTD) with genetically defined subgroups, the sparse GFA identified latent disease factors differentially expressed across the subgroups, distinguishing between "subgroup-specific" latent factors within homogeneous groups and "subgroup common" latent factors shared across subgroups. The latent disease factors captured associations between brain structure and non-imaging variables (i.e., questionnaires assessing behaviour and disease severity) across the different genetic subgroups, offering insights into disease profiles. Importantly, two latent factors were more pronounced in the two more homogeneous FTD patient subgroups (progranulin (GRN) and microtubule-associated protein tau (MAPT) mutation), showcasing the method's ability to reveal subgroup-specific characteristics. These findings underscore the potential of sparse GFA for integrating multiple data modalities and identifying interpretable latent disease factors that can improve the characterization and stratification of patients with neurological and mental health disorders.

因子分析疾病分型多模态数据神经疾病

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。