arXiv:2604.24796q-bio.OTcs.LG2026-04

用多阶段计算框架解析肝硬化数据,提升疾病建模与治疗决策精度。

A multi-stage soft computing framework for complex disease modelling and decision support: A liver cirrhosis case study

论文配图:A multi-stage soft computing framework for complex disease modelling and decision support: A liver cirrhosis case study
图 1 · 摘自论文原文
  • 融合单细胞测序与基因网络分析,稳定高维噪声数据中的基因模块。
  • 用卷积神经网络构建疾病二维图谱,分类性能优于传统方法。
  • 可推广至其他复杂疾病,适合需解释性与小样本建模的研究者。

肝硬化是全球重大健康问题,每年导致数百万死亡,早期检测与积极治疗可显著改善患者生活质量。从生物医学数据建模复杂疾病面临高维、强特征相关性、噪声及标注样本少等挑战,传统机器学习方法在鲁棒性、可解释性和泛化能力上常显不足。本研究提出一种面向复杂疾病建模与治疗探索的多阶段机器学习决策框架。该框架整合单细胞转录组分析、高维加权基因共表达网络(hdWGCNA)特征稳定、多模型学习、深度表征构建及事后决策支持。具体而言,通过单细胞测序识别关键细胞亚群,利用hdWGCNA在稀疏与噪声条件下稳定基因模块;为增强非线性特征交互建模,将表格型分子特征重构为二维疾病图谱并使用卷积神经网络(CNN)分析;最后引入分子对接作为决策支持模块评估候选药物。以肝硬化为例,框架识别出与疾病相关的内皮细胞亚群,并提取出7个稳健标志基因(HSPB1, GADD45A, CLDN5, ATP1B3, C1QBP, ENPP2, PARL)。基于CNN的表征学习模块在分类任务中表现优于传统流程。该框架具有疾病无关性,可直接扩展至其他涉及不确定性、异质性和小样本的组学驱动生物医学应用。

原文摘要 · Abstract (English)

Liver cirrhosis is a major global health problem causing millions of deaths annually, and timely detection with aggressive treatment can significantly improve patients' quality of life. Modelling complex diseases from biomedical data is computationally challenging due to high dimensionality, strong feature correlations, noise, and limited labelled samples. Conventional Machine Learning (ML) pipelines often struggle with robustness, interpretability, and generalisation under such conditions. In this study, we propose an ML-driven multi-stage decision framework for complex disease modelling and therapeutic exploration. The framework integrates single-cell transcriptomic profiling, high-dimensional network-based feature stabilisation, multi-model learning, deep representation construction, and post-hoc decision support. Specifically, single-cell sequencing data were analysed to identify key cellular subpopulations, followed by high-dimensional weighted gene co-expression network analysis (hdWGCNA) to stabilise gene modules under sparsity and noise. To enhance non-linear feature interaction modelling, tabular molecular features were restructured into two-dimensional disease maps and analysed using a CNN. Finally, molecular docking was incorporated as a decision-support module to evaluate candidate therapeutic compounds. Using liver cirrhosis as a representative case, the framework identified a disease-associated endothelial subpopulation and extracted seven robust signature genes (HSPB1, GADD45A, CLDN5, ATP1B3, C1QBP, ENPP2, and PARL). The CNN-based representation learning module outperformed conventional pipelines in classification. The framework is disease-agnostic and readily extends to other omics-driven biomedical applications involving uncertainty, heterogeneity, and limited samples.

疾病建模单细胞测序深度学习肝硬化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。