提出新方法GGRF,高效分析测序数据中的罕见基因变异与疾病关联。
A Generalized Genetic Random Field Method for the Genetic Association Analysis of Sequencing Data
- 基于广义估计方程框架,无需设定罕见变异阈值。
- 在多种疾病模型下,检测力优于或不逊于SKAT,尤其对罕见变异敏感。
- 适用于小样本数据,适合研究罕见变异致病机制的科研人员。
随着高通量测序技术的发展,全面分析序列变异对复杂人类疾病的影响已成为可能。尽管利用新测序技术的研究有望揭示新的遗传变异,特别是贡献于疾病的罕见变异,但高维测序数据的统计分析仍具挑战性。本文提出一种广义遗传随机场(GGRF)方法,用于测序数据的关联分析。该方法继承了相似性分析方法(如SIMreg和SKAT)的优势,无需设定罕见变异的阈值,可检验多变异以不同方向和效应大小作用的情况。方法基于广义估计方程框架,支持多种疾病表型(如定量和二分类)。此外,具有良好的渐近性质,可在小样本数据中直接应用而无需小样本校正。模拟结果显示,在多种疾病情景下,尤其是罕见变异在致病中起关键作用时,GGRF的检验力优于或接近常用方法SKAT。进一步在达拉斯心脏研究的真实数据中应用,成功发现候选基因ANGPTL3和ANGPTL4与血清甘油三酯水平相关。
原文摘要 · Abstract (English)
With the advance of high-throughput sequencing technologies, it has become feasible to investigate the influence of the entire spectrum of sequencing variations on complex human diseases. Although association studies utilizing the new sequencing technologies hold great promise to unravel novel genetic variants, especially rare genetic variants that contribute to human diseases, the statistical analysis of high-dimensional sequencing data remains a challenge. Advanced analytical methods are in great need to facilitate high-dimensional sequencing data analyses. In this article, we propose a generalized genetic random field (GGRF) method for association analyses of sequencing data. Like other similarity-based methods (e.g., SIMreg and SKAT), the new method has the advantages of avoiding the need to specify thresholds for rare variants and allowing for testing multiple variants acting in different directions and magnitude of effects. The method is built on the generalized estimating equation framework and thus accommodates a variety of disease phenotypes (e.g., quantitative and binary phenotypes). Moreover, it has a nice asymptotic property, and can be applied to small-scale sequencing data without need for small-sample adjustment. Through simulations, we demonstrate that the proposed GGRF attains an improved or comparable power over a commonly used method, SKAT, under various disease scenarios, especially when rare variants play a significant role in disease etiology. We further illustrate GGRF with an application to a real dataset from the Dallas Heart Study. By using GGRF, we were able to detect the association of two candidate genes, ANGPTL3 and ANGPTL4, with serum triglyceride.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。