提出新算法,从临床数据中高效识别同质患者亚群。
A Goemans-Williamson type algorithm for identifying subcohorts in clinical trials
- 基于广义的Goemans-Williamson思想设计舍入策略。
- 保证近似最优解,逼近比达0.82倍。
- 应用于乳腺癌数据,发现关键基因关联线索。
我们设计了一种高效算法,用于从大规模异质数据集中识别出主要同质的患者亚群。理论贡献在于一种类似Goemans与Williamson(1995)的舍入技术,可将最优解近似至0.82倍。作为应用,我们利用该算法在Curtis等人(2012)的乳腺癌RNA微阵列数据集中系统性地权衡敏感性与特异性,识别出具有临床意义的同质亚群。其中一个亚群提示:LXR过表达与BRCA2和MSH6甲基化水平存在关联。
原文摘要 · Abstract (English)
We design an efficient algorithm that outputs tests for identifying predominantly homogeneous subcohorts of patients from large in-homogeneous datasets. Our theoretical contribution is a rounding technique, similar to that of Goemans and Wiliamson (1995), that approximates the optimal solution within a factor of $0.82$. As an application, we use our algorithm to trade-off sensitivity for specificity to systematically identify clinically interesting homogeneous subcohorts of patients in the RNA microarray dataset for breast cancer from Curtis et al. (2012). One such clinically interesting subcohort suggests a link between LXR over-expression and BRCA2 and MSH6 methylation levels for patients in that subcohort.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。