当前位置: 首页>博士论文>资源详情
面向复杂性状遗传性缺失的关联分析方法研究
中文摘要

 复杂性状作为一种受多个微效基因和环境共同控制的性状,其遗传机理研究一直是遗传学的热点与难点。复杂性状的研究在人类复杂疾病的预防、诊断和治疗、以及改良作物农艺和品质性状、选育优良品种中都发挥着至关重要的作用。随着人类及常见动植物基因组测序工作的相继完成,全基因组关联分析方法(genome-wide association studies, GWAS)成为当下研究基因与复杂性状关联性最重要的策略与工具之一。但对大多数复杂性状而言,由全基因组关联分析确定的关联位点只能够解释相应疾病遗传方差的5%到30%,这一现象被称为“遗传性缺失”。已有研究表明,导致该现象发生的主要因素包括:1)单位点分析统计效力不足;2)缺乏对基因之间相互作用以及基因-环境相互作用的分析;3)罕见变异对性状的影响尚不明晰。基于此,本文以单核苷酸多态性(single nucleotide polymorphism, SNP)数据为数据材料,从上述1)和2)两个关键因素出发,对多位点关联分析、上位效应检测以及基因-基因相互作用等计算问题进行深入研究,为解决“遗传性缺失”提供新的思路和方法。具体内容包括以下四个方面: (1)提出基于混合线性模型和稀疏组Lasso的多位点关联分析方法 为解决传统GWAS中单位点关联分析统计效力不足的问题,提出一种基于混合线性模型和稀疏组Lasso的多位点关联分析方法。首先,针对传统单一位点检测方法中存在的多检验校正导致的大量位点无法通过严格阈值、以及无法利用位点之间的关联性两个弊端。将多元线性模型引入关联分析,通过同时建模多个位点与表型之间的关联性,充分利用多个位点的联合信息来提升方法统计效力。其次,针对数据中存在的种群结构等混淆因素导致方法假阳性升高的问题,通过在线性模型中显式将混淆因素建模为随机效应项,更有针对性地消除系统误差对方法精度的影响。最后,通过使用稀疏组Lasso方法优化求解模型,使模型具有基因层面与SNP层面的双重稀疏性。实验结果表明,该方法在可以有效降低混淆因素造成的假阳性,提升在性状预测及关联位点选取方面的准确性,成为有力的关联分析工具。 (2)提出基于因子分解机的上位效应关联分析方法 从SNP数据层面研究变异位点之间的相互作用关系(即上位效应)被认为是解决“遗传性缺失”问题的可行性方案之一。现有方法中,基于穷举的上位效应检测策略导致计算量随着作用关系涉及位点数目的增加呈幂级增长。而基于随机或启发式的检测方法会因优化目标的不同导致上位效应的丢失。如何在考虑所有位点组合的情况下降低方法的计算复杂度,是上位效应检测的一个发展方向。基于此,本文提出一种基于因子分解机的上位效应检测方法。首先,通过独热编码对基因型数据进行稀疏化,获得等位基因层面对上位效应遗传机理的解释;其次,利用因子分解机学习每个向量在隐空间的嵌入向量表示,并利用两个特征之间嵌入向量的内积表征上位效应的作用强度。因子分解机可以在稀疏数据上有效学习交互特征的作用关系,在线性时间复杂度下完成2阶上位效应检测。实验结果表明,本文提出的基于因子分解机的上位效应检测方法可以高效、准确地检测互作关系。 (3)提出质量性状下基因-基因相互作用的关联分析方法 变异位点层面的相互作用研究往往会带来组合爆炸、统计效力低等问题。近年来,基于基因整体(即将一个基因中的所有SNP看做一个整体)的基因互作研究成为GWAS中的又一热点内容。本文提出一种基于距离相关系数和置换检验策略的基因-基因相互作用检验方法。首先,利用距离相关系数对非线性交互作用较强的检测能力,以及对两个交互向量维度不设限的优势,构造表示2组SNP在疾病样本与对照样本中相互作用差异的统计量来表征基因相互作用关系的强度。所设计统计量对基因之间的互作形式没有限制,可以使方法具有更好的泛化能力;其次,针对本文所设计统计量经验分布未知的问题,利用置换检验策略近似其分布,从而得到基因互作强度的显著性指标。实验结果表明,本文所提出的方法显著优于其它方法,可以有效准确地检测基因-基因之间不同形式的相互作用关系。 (4)提出数量性状下基因-基因相互作用的关联分析方法 数量性状取值在群体内个体间呈现连续性。研究与人体内脂质水平相关的遗传变异对于理解心脑血管疾病的致病机理至关重要;而植物的开花时长、粒重等也与优良品种的选育息息相关。但目前的研究中,针对数量性状下基因互作的研究十分有限。因此,提出一种基于U统计量和集成学习的检验方法,用于检验数量性状下基因的交互作用。首先,针对互作中大量存在的非线性关系,选取集成学习模型作为学习算法,其中集成学习的基分类器选用树模型,充分捕捉数据中不同形式的作用关系的同时保证模型的泛化能力;其次,通过采用使预测结果具有U统计量渐近正态性质的重采样策略,设计用于表征互作关系强度的统计量。实验结果表明,本文提出的方法可以有效检测数量性状下基因-基因之间不同形式的相互作用关系。 关键词:全基因组关联分析;单核苷酸多态性位点;遗传性缺失;多位点关联分析;上位效应;基因-基因相互作用

英文摘要

 Complex traits, being controlled jointly by multiple genes as well as environmental factors, have always been an important and challenging research topic in genetics. The study of these traits has important applications in prevention, diagnosis, and treatment of complex human diseases as well as in the selective breeding of crops. Since the completion of the sequencing of human genomes as well as the genomes of many common plants and animals, genome-wide association studies (GWAS) have been successfully applied to the study of many complex phenotypical traits/diseases, and has become one of the most important tool in the study of the relationship between genes and complex traits. However, despite the major success in uncovering controlling genes of complex human diseases, for most complex traits GWAS can only detect genetic variants that account for between 5% to 30% of the variance attributed to genetic factors, which is called the "missing heritability problem". Based on current research, some potential causes of this problem includes 1) insufficient statistical power of single locus analysis; 2) lack of considerations of gene-gene and gene-environment interactions; 3) insufficient study of the effect of rare genetic variants. Hence, using single nucleotide polymorphism (SNP) data, we will focus on the issues 1) and 2) above and propose four new strategies for multi-locus modeling as well as the modeling of epistasis and gene-gene interactions: (1)Multi-locus association analysis via combining linear mixed model and sparse group Lasso In order to deal with the deficiency in the statistical power of traditional single-locus GWAS, we proposed a method for multi-locus association studies based on linear mixed model and sparse group lasso. Firstly, in order to handle the high threshold due to multi-test correction which is unavoidable in single locus GWAS, as well as the inability of utilizing connections between loci, we used a multivariant linear model in association studies in order to fully utilize the combined information from multiple loci and strengthen the statistical power. Secondly, to deal with the potential false positives caused by confounding factors like population structure, we used the linear mixed model and model confounding factors via random effects. Lastly, based on the biological intuition that phenotypes should be related to a few SNPs in a few genes only, we incorporate prior knowledge on the location of the SNPs via sparse group lasso. Experiments in both simulated and real data show that our new multi-locus method performs well in both phenotype prediction and effective locus selection, and is a powerful tool for analysis studies. (2)Epistasis detection based on factorization machines The study of interactions between SNPs (epistasis) is an important topic in GWAS. Among the current method, statistically testing the interactions of all pairs of SNPs has a time complexity that grows quadratically with the number of SNPs and may have insufficient statistical power due to the use of multi-test correction, while randomized or heuristic search methods may not be able to find all epistasis. Hence, a major problem in the study of epistasis is how to reduce complexity while still considering the interaction of all possible combinations of SNPs. In this article we proposed an epistasis detection method based on factorization machines: firstly, do one-hot coding for SNP data to turn them into sparse features, then use these sparse data as input of the factorization machine to learn embedding vectors, then use the inner product of embedding vectors to represent the strength of interactions between pairs of SNPs. Our experiments show that compared with prior approaches, this method can detect epistasis more accurately and efficiently. (3)Gene-based nonparametric testing of interaction underlying qualitative traits As marker-level interaction studies may have high time complexity and low statistical power, recently interaction between genes (i.e. groups of SNPs in the same gene seen as one single feature) become another popular topic in GWAS. Here we propose a new strategy for the detection of gene-gene interactions based on a permutation strategy and distance correlation (dcor). Firstly, due to the fact that dcor has strong power for detecting nonlinear interaction, and has no constraint on the dimension of the two features, we use the difference of dcor in disease and control samples to characterize the strength of the interaction. Because there is little assumption on the exact form of interaction when measuring it via the difference of dcor, this method should have good performance under generalization. Furthermore, to get the significance of the dcor, we use a permutation strategy to estimate the distribution of dcor in the absence of interactions. Experiments on 8 simulated disease models as well as the real data on human rheumatoid arthritis (RA) shows that it has significant advantages over prior approaches for the detection of gene interactions. (4)Gene-based testing of interaction underlying quantitative traits As opposed to qualitative traits, quantitative traits are those that take continuous values among the individuals in a group. The lipid level in humans, the flowering time in plants, and the weight of grains are all belong to quantitative traits. Investigating the genetic variation related to lipid levels in the human is crucial for understanding the pathogenesis of cardiovascular and cerebrovascular diseases; and the flowering time and grain weight of plants are also closely related to the breeding of elite varieties. However, currently, researches on genetic interaction underlying quantitative traits are limited. Therefore, a test method based on U statistic and ensemble learning method was proposed to test the interaction of genes underlying quantitative traits. Firstly, aiming to model the nonlinear relationship in the gene-gene interaction, we introduce the ensemble learning method as our learning algorithm. To guarantee the ability to fully capture the different forms of interaction as well as good generalization performance, we choose tree model as the base learner. Using regression tree as base learner. Secondly, a special subsampling way is used to ensure the prediction made by ensemble method belong to U-statistics that we could use its approximal normal distribution property to design statistics that modeling the strength of gene-gene interaction. Experimental results show that the proposed statistical test can effectively detect different forms of the gene-gene interaction underlying quantitative traits. Keywords: Genome-wide association study, Single nucleotide polymorphism, missing heritability, multi-locus association analysis, epistasis, gene-gene interaction

作者相关
主题相关
看过该书的人还在看哪些书