摘要:为了加深对人类疾病的致病机制的理解、提高其诊断和治疗效率,疾病相关的作用关系研究一直都是生物信息中研究的热点问题。研究发现疾病不但与个人生活习惯和环境因素有关,也与遗传存在紧密的关联关系,特别是癌症类的复杂疾病。当前,研究人员进一步发现不仅仅基因是疾病的致病原因,其它的生物分子与疾病也有着密不可分的关系,比如miRNA和微生物等。同时,为了提高疾病的治疗效率和降低成本,作为药物开发和疾病治疗中必不可少的药物-靶标关系和药物-药物关系研究也是当前疾病关联关系研究的重要方向。本文从生物信息学的角度出发,对疾病致病机制、诊断和治疗相关的关联关系问题展开了研究。包括疾病发病机制和生物标记相关的miRNA-疾病关系预测和微生物-疾病关系预测,以及疾病治疗相关的药物-靶标关系预测和药物-药物关系预测。在对上述问题当前研究状况进行理解和分析基础上,结合相关的生物分子网络信息,利用计算模型对预测问题进行求解。全文的主要研究创新点如下: (1)基于已知的miRNA-疾病关联关系已经得到实验验证的特点,针对当前预测方法对已知关系与未知关系赋予同等权重的问题,本文提出了一种集成miRNA和疾病相似性以及对己知关系赋予更高权重的方法(DNRLMF-MDA)来预测miRNA-疾病关系。DNRLMF-MDA集成了已知的miRNA疾病关系,miRNAs的功能相似性和高斯核相似性,疾病功能相似性和高斯核相似性。特别的是,对己知确认的miRNA疾病关联关系比未知的关联关系赋予了更高的权重。DNRLMF-MDA通过逻辑矩阵分解方法来计算一个miRNA和疾病对的关联关系分数,潜在的miRNA和疾病向量表示miRNA和疾病的属性,通过动态邻居正则化进一步提高了其预测性能。5-倍交叉验证用来验证DNRLMF-MDA的预测性能并与其他方法进行比较。实验结果表明,DNRLMF-MDA优于PBMDA,HGBI和NBI三个方法。 (2)提出了一种基于相似性和双随机游走的微生物-疾病预测方法(BRWMDA)。在本研究中,BRWMDA将微生物网络,疾病网络和已知的微生物-疾病关联关系整合到一个异构网络中,通过随机游走方法对潜在的微生物-疾病关系进行预测。在构建微生物网络过程中,对微生物的高斯核相似性通过逻辑函数进行了调节处理。疾病网络的创建通过对疾病的高斯核相似性和疾病语义相似性进行融合集成得到。基于相似的微生物通常关联相似的疾病(反正亦然)的假设,BRWMDA采用改进的双随机模型来预测潜在的微生物-疾病关联关系。通过5-倍交叉验证和留一交叉验证进行性能评估并与其他方法进行比较。实验结果表明,BRWMDA的预测性能优于 NGRHMDA,LRLSHMDA,KATZHMDA,HGBI和NBI五个方法。 (3)当前的药物-靶标关系预测方法主要针对存在已知关系的药物进行预测,而对失败的药物和全新的化学实体的大规模预测的方法很少,但是此类数据广泛存在于当前的一些生物数据库。在本文中提出了一种基于药物化学子结构的药物-靶标关系预测方法(SDTRLS)。SDTRLS融合药物化学子结构相似性和已知药物靶标关系的高斯核相似性,通过克罗内克积最小二乘法(RLS-Kron)模型预测最终的药物靶标关系。SDTRLS通过大规模生物分子网络信息为已有的药物,失败药物和新化学实体物质进行药物关系预测。10-倍交叉验证和外部验证用来评估算法的预测性能并与其他方法进行比较。SDTRLS能够对存在己知药物-靶标关系的药物,失败的药物和新的化学实体的新的药物靶标关系进行预测。通过实验表明本文提出的SDTRLS的预测性能优于最新的SDTNBI方法。 (4)针对当前对已知对药物相关分子信息在药物-药物关系预测中的利用不够和全新的药物的关联关系预测效果不明显的问题,提出了一种基于药物相似性集成的药物-药物关系预测方法(DDIRLS)。在本研究中, DDIRLS集成了药物的化学,生物,表型方面的信息和己知的药物-药物关系。通过基于药物化学,生物学,表型信息的余弦相似性和药物的高斯核的均值集成来得到药物相似性,利用最小二乘法来进行药物-药物关系预测。在对全新的药物的关联关系预测过程中,通过基于结点的网络扩散的方式来初始化其关联关系分数,进一步提高了其预测性能。利用5-倍交叉验证,10-倍交叉验证和denovo药物验证,本文系统的评估了本算法的预测性能并与其他方法进行了比较。实验结果表明,DDIRLS的预测性能优于WAE,L1E,L2E方法。 图20幅,表28个,参考文献247篇。 关键词:疾病;MicroRNA;微生物;药物-药物关系;药物-靶标关系;机器学习 分类号:TP301.6;Q71
Abstract: In order to deepen the systematic understanding of the pathogenesis of human diseases, improve their diagnosis and treatment efficiency, disease-related relationship research has always been a hot issue in bioinformatics. Many studies found that the disease is not only related to personal living habits and environmental factors, but also closely related to heredity, especially complex diseases of cancer. Currently, researchers have further discovered that not only genes are the cause of disease, but other biomolecules are also inextricably linked to diseases such as miRNAs and microorganisms. At the same time, in order to improve the treatment efficiency and reduce the cost of the disease, as the drug-target interaction and drug-drug interaction research are essential for drug development and disease treatment, they are also the important direction of current disease relationship research. From the point of view of bioinformatics, this paper studies the disease-related associations. These include miRNA-disease interaction prediction and microbe-disease association prediction related disease pathogenesis, drug-target interaction prediction and drug-drug interaction prediction related disease treatment. Based on the systematically understanding and analyzing of the current progress of the above problems, the prediction problem is solved by the computational model via integrating the relevant biomolecular network information. The main study and innovations of this paper are in follows: (1)Based on known miRNA-disease associations are validated by biological experiment and current methods for miRNA-disease association prediction assign the same important level for known associations and unknown associations, we propose a method (DNRLMF-MDA) to predict microbe-disease association, which integrates the miRNA and disease similarity and assignes the higher weight value to known associations than unknowns. DNRLMF-MDA integrates known miRNA-disease associations, functional similarity and Gaussian Interaction Profile (GIP) kernel similarity of miRNAs, and functional similarity and GIP kernel similarity of diseases. Especially, positive observations (known miRNA-disease associations) are assigned higher importance levels than negative observations (unknown miRNA-disease associations). DNRLMF-MDA computes the probability that a miRNA would interact with a disease by a logistic matrix factorization method, where latent vectors of miRNAs and diseases represent the properties of miRNAs and diseases, respectively, and further improve prediction performance via dynamic neighborhood regularized method. The 5-fold cross validation is adopted to assess the performance of our DNRLMF-MDA, as well as other competing methods for comparison. The computational experiments show that DNRLMF-MDA outperforms other three (PBMDA, HGBI and NBI) methods. (2) A method (BRWMDA) is proposed to predict new microbe-disease associations based on similarity and bi-random walk on the disease and microbe networks. In this study, BRWMDA integrates the microbe network, disease network and known microbe-disease associations into a heterogeneous network, and predicts potetional microbe-disease associations via bi-random walk on this network. In constructing microbe network, BRWMDA adds the adjustment process of microbe GIP similarity by logistic function. Then constructing the diseases network by SNF mothed based on GIP similarity and symptom-based similarity of diseases. Based on assumption that similar microbes are normally associated with similar diseases and vice versa, BRWMDA is adopted to predict new potential microbe-disease associations via using bi-random walk model. The 5-fold cross validation and Leave One Out Cross Validation (LOOCV) are adopted to assess the prediction performance of our BRWMDA algorithm, as well as other competing methods for comparison. The experiment results show that the prediction performance of BRWMDA outperforms other five methods (NGRHMDA, LRLSHMDA, KATZHMDA, HGBI and NBI). (3) Existing computational methods are mainly deal with drug-target interaction (DTI) predictions of known drugs, there are few methods for large-scale prediction of failed drugs and new chemical entities that are currently widely existed in some biological databases. In this study, a method (SDTRLS) is proposed to predict drug-target interaction based on chemical substructure of drugs. SDTRLS adopt RLS-Kron model to predict DTIs by integrating GIP similarity and chemical substructure similarity of drugs. The 10-fold cross validation and External validation are used to evaluate the prediction performance of SDTRLS, as well as other competing methods for comparison. SDTRLS can be an effective predictor for targets of old drugs, failed drugs and new chemical entities by using large scale biomolecular network databases. The computational experiments results show that SDTRLS outperforms the state-of-art SDTNBI method. (4) Because the current computational methods are not enough to effectively use the molecular networks information of drugs, and also ignore the drug-drug interactions (DDIs) prediction for new drugs, a new method (DDIRLS) is proposed to predict the potential DDIs based on integrated similarity of drugs. In this study, DDIRLS integrating the chemical, biological, phenotype data and known DDIs. Based on the drug feature similarity by cosine method via integrating drug chemical, biological, phenotype data and drug GIP similarity with known DDIs, we obatine the drug final similarity by averaging. DDIRLS adopt Regularized Least Squares (RLS) classifier to predict DDIs. The interaction scores of new drugs are initialized by the node-based drug networks diffusion method, which further improves the prediction perfromance. By using the 5-fold cross validation, 10-fold cross validation and denovo drug validation, we systematically assess the prediction performances of DDIRLS and compare it with other prediction methods. The experiment results show that DDIRLS outperforms the WAE (Weighted average ensemble method), L1E (Classifier ensemble method (L1)), L2E (Classifier ensemble method (L2)) methods. Keywords: Disease; MicroRNA; Microbe; Drug-drug interaction; Drug-target interaction; Machine learning Classification: TP301.6;Q71