数据特性比算法更重要,决定脑影像分型能否成功
Dataset Properties Shape the Success of Neuroimaging-Based Patient Stratification: A Benchmarking Analysis Across Clustering Algorithms
- 用合成数据测试四种算法,发现数据复杂度影响远超算法选择
- 集群重叠、大小不均或效应微弱时,准确率最高下降50%
- 适合关注脑影像分型的科研人员,尤其需重视数据预处理
背景:基于神经影像的数据驱动患者分型有望推动精准神经精神医学发展,但现有聚类方法常无法跨队列泛化。尽管算法创新多聚焦模型复杂度,但数据集本身特性仍被忽视。我们假设输入数据中的簇分离度、大小不平衡、噪声水平及疾病效应的方向与强度,是决定聚类准确性和可重复性的关键因素。方法:在基于人类连接组计划青年成人数据集生成的合成脑形态学队列上,评估了四种常用分型算法(HYDRA、SuStaIn、SmileGAN、SurrealGAN)。对600名伪患者施加三种全局变换模式,对比508名对照,再通过4种内部变体调整簇数(k=2–6)、重叠度和效应量。性能以恢复已知真实簇的准确率衡量。结果:在122种合成场景中,数据复杂度始终比算法选择更能预测分型成功。簇间清晰分离时,所有方法准确率均高;而簇重叠、大小不均或效应微弱时,准确率最高下降50%。SuStaIn无法扩展至17个以上特征,HYDRA在数据异质性下表现不稳定。SmileGAN与SurrealGAN虽能稳健检测模式,但未为个体分配离散簇标签。结论:研究揭示了输入数据统计特性对各类算法的普遍影响,强调新算法开发应使用真实数据分布,并建议更多采用以数据为中心的策略,主动塑造和标准化输入分布。
原文摘要 · Abstract (English)
Background: Data driven stratification of patients into biologically informed subtypes holds promise for precision neuropsychiatry, yet neuroimaging-based clustering methods often fail to generalize across cohorts. While algorithmic innovations have focused on model complexity, the role of underlying dataset characteristics remains underexplored. We hypothesized that cluster separation, size imbalance, noise, and the direction and magnitude of disease-related effects in the input data critically determine both within-algorithm accuracy and reproducibility. Methods: We evaluated 4 widely used stratification algorithms, HYDRA, SuStaIn, SmileGAN, and SurrealGAN, on a suite of synthetic brain-morphometry cohorts derived from the Human Connectome Project Young Adult dataset. Three global transformation patterns were applied to 600 pseudo-patients against 508 controls, followed by 4 within-dataset variations varying cluster count (k=2-6), overlap, and effect magnitude. Algorithm performance was quantified by accuracy in recovering the known ground-truth clusters. Results: Across 122 synthetic scenarios, data complexity consistently outweighed algorithm choice in predicting stratification success. Well-separated clusters yielded high accuracy for all methods, whereas overlapping, unequal-sized, or subtle effects reduced accuracy by up to 50%. SuStaIn could not scale beyond 17 features, HYDRA's accuracy varied unpredictably with data heterogeneity. SmileGAN and SurrealGAN maintained robust pattern detection but did not assign discrete cluster labels to individuals. Conclusions: The study results demonstrate the impact of statistical properties of input data across algorithms and highlight the need for using realistic dataset distributions when new algorithms are being developed and suggest greater focus on data-centric strategies that actively shape and standardize the input distributions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。