人类胚胎干细胞是指胚泡期的内细胞团中分离出的多潜能细胞。胚胎干细胞具有许多特异的生物学特性:①全能性。在去除分化抑制的条件下,胚胎干细胞能发育为构成机体任何一种细胞的潜力。②自我更新,无限增殖性。胚胎干细胞在适宜条件下,能长久而稳定地自我更新增殖,达到“永生”。由于其具有以上独特的性质,ES细胞在再生医学、药理学及发育学等领域具有重要应用前景。因此,胚胎干细胞研究吸引着世界各国许多科学家的关注,目前已经鉴定出了一些胚胎干细胞重要基因,如OCT3/4,NANOG,REX1,SOX2和FOXD3等,同时对部分基因、蛋白进行了较为详尽的分子生物学研究。分子生物学,主要采用还原论式思维方式,它通过解析生物系统中特定基因或者蛋白质,从而推演生命现象的物质基础和生命过程的基本活动规律。分子生物学帮助我们在分子水平上认识了生物分子的结构和功能,促进了生物医学的发展。但是,分子生物学研究仅仅强调生物体内部单个分子,未考虑分子之间相互作用,对解析复杂系统存在较大缺陷。众所周知,人体内的蛋白和其他一些小分子不是单独起作用,而是相互作用形成一种分子相互作用网络,这种相互作用网络决定了细胞乃至组织、个体的特征。直到目前为止,大多关于干细胞的科学研究仅仅强调单个基因和蛋白,没有考虑他们之间可能联系和作用。关于胚胎干细胞特征的了解仍然非常有限,这也限制了它的临床应用。 随着后基因组时代的到来,高通量研究技术取得了巨大发展,如基因芯片技术,蛋白芯片技术等,因此生命科学研究形式也开始发生转变,如由过去一味强调单基因、蛋白的研究开始转向研究整个相互作用组。目前,生物医学知识常常可以采用网络形式来代表,如调节网络、代谢网络、基因相互作用网络、蛋白相互作用网络、小分子相互作用网络等。构建和分析这些网络揭示了许多以前未知的知识。在人类胚胎干细胞研究中,网络知识也得到广泛的应用,但大多研究强调转录调控网络研究,在此种网络研究中,主要探讨转录因子的重要性,这些研究发现了一些在人类胚胎干细胞特性调节中许多重要的转录因子,最近,基于蛋白-蛋白相互作用的网络生物信息学研究取得了较大成功。在胚胎干细胞蛋白相互作用网络研究中,采用NANOG为诱饵的亲和纯化加上质谱研究产生了由37个老鼠胚胎干细胞蛋白组成的相互作用网络。虽然此项研究取得了较好的开端,但涉及的蛋白数较少,研究并不完善。 芯片技术提供了一个用于同时研究一个细胞或组织全基因组表达的较好平台。在胚胎干细胞基因表达研究中,许多实验室采用芯片技术,但由于各种原因,如芯片平台差异、细胞株差异、实验操作技术差异等,人类胚胎干细胞基因表达谱存在较大的异质性。最近采用统合分析方法分析基因表达谱芯片增加了结果的可靠性。Assou等采用此方法得到了一组“人类胚胎干细胞一致基因集”。在此基因集中的一些基因同时也在人体其他少数组织细胞中表达,说明一些蛋白单个形式表达可能不是特异性存在于胚胎干细胞,但是整个联合蛋白集的表达形式是胚胎干细胞特异的,这暗示它们之间的相互作用可能是特异性的。我们推测,胚胎干细胞可能存在一个富集蛋白相互作用网络调节和维持它们的特性。 复杂网路无标度特性的提出,突破了随机网络模型的束缚,使大家认识到各种复杂系统的网络结构,都遵从某些基本法则。随后几年,全世界范围内兴起了研究复杂网络的热潮。复杂网络理论以社会网络、技术网络、生物网络等真实网络为研究对象,通过图论等方法,研究网络结构特征、结构与功能的关系等一系列问题,从而来获得对于现实系统更多认识。复杂网络理论作为一门新兴学科,为在系统水平上研究生物网络提供了新的理论依据和平台。2000年Jeong等人在Nature上第一次发表利用复杂网络理论研究代谢网络拓扑特性的论文,自此以后,利用复杂网络理论研究各种生物网络迅速发展。复杂网络理论揭示,细胞内分子相互作用网络的结构特性,与其它复杂系统网络(如万维网、社会网)在很大程度上是一致的,说明可能存在相似的法则控制着多数现实中的生物复杂网络系统。 根据以上研究,我们提出两个问题:(1)是否存在胚胎干细胞富集的蛋白相互作用网络?如果有,此网络的拓扑结构特征怎样?(2)是否存在胚胎干细胞富集的功能相互作用方式,调节胚胎干细胞特性?我们采用基于复杂网络理论的生物信息学方法来解释以上问题。 研究内容主要分为三个部分: 第一部分:构建人类胚胎干细胞富集的蛋白相互作用网络,并探讨网络拓扑特征。收集已有研究报道采用统合分析得到的一组“人类胚胎干细胞一致基因集”,通过在线Uniprot ID软件进行名称转换,得到1020个UniprotKb/SwissProt蛋白号,即胚胎干细胞相关蛋白。同时下载人类蛋白相互作用数据库I2d,经自编perl程序对数据库进行整理,删除数据库中冗余的蛋白相互作用对,得到13560个SwissProt蛋白及其组成的92545个非冗余蛋白相互作用对。收集了课题组梁爽教授研究报道97个正常人类组织中选择性表达的基因Affymetrix探针号,通过Affymetrix公司提供的最新注解文件对探针号进行重新注解,最终获得3904个组织选择性编码蛋白基因。对比干细胞一致性基因集和组织选择性基因,发现有274交叉基因。利于自编perl程序寻找由胚胎干细胞相关蛋白组成的蛋白相互作用对,通过广度优先算法搜索I2d蛋白相互作用数据库,得到由403个胚胎干细胞相关蛋白组成的连续蛋白相互作用网络。并进行1000次随机抽样网络,统计分析表明明显小于此干细胞网络,我们将之命名为干细胞富集蛋白相互作用网络,通过Cytoscape软件对此网络进行可视化分析。根据复杂无标度网络Barabasi-Albert模型,分析构建的干细胞富集蛋白相互作用网络中各节点度及其相关度分布,最终得到网络度相关幂律指数γ值为1.3081,证明此网络和真实网络类似,具有无标度特性。 第二部分:根据复杂网络理论,重点对网络中心蛋白进行分析。由于干细胞富集蛋白相互作用网络具有无标度特性,我们采用自编perl程序,利用Dijkstra算法对不同删除方法删除网络一定节点后,计算网络平均最短路径长度变化,探讨网络是否具有鲁棒性和脆弱性特征,结果发现当删除网络0、4、8、 12、16、20个随机节点后,网络的平均最短路径长度分别为:3.679、3.676±0.006、 3.674±0.016、3.672±0.016、3.686±0.052、3.688±0.040。经统计证明删除不同数量节点后与未删除节点的平均最短路径长度未发生明显改变。当删除网络0、4、 8、12、16、20个中心节点后,网络平均最短路径长度分别为3.679、3.770±0.055、 3.849±0.065、4.028±0.020、4.208±0.118、4.448±0.092。经统计证明删除不同中心节点与未删除节点比较,平均最短路径长度发生了明显改变,而且随着删除中心节点数目增加,平均最短路径长度随之增加。以上证明我们构建的网络具有较强的鲁棒性和脆弱性,说明网络中心蛋白(节点)对网络的拓扑结构稳定具有重要作用。采用5%最高连接数标准定义中心蛋白,我们共发现21个中心蛋白,分别为:MYC、EIF4A1、DDX18、H2AFX、KIAA0020、RPL4、PCNA、 POLR1B、HSPA8、DKC1、CDC2、EIF4E、BXDC1、BOP1、RPLP0、EIF3A、 RUVBL1、HDAC2、GFPT2、HIST1H4C、CBS。通过文献搜索,我们发现中心蛋白中MYC、H2AFX、RUVBL1已有报导与干细胞自我更新,增殖等特性密切相关。另外通过MGI数据库资料发现POLR1B、CDC2、HDAC2、MYC与胚胎发育密切相关,突变将导致胚胎致死,结合上述网络研究结果,我们推测:中心蛋白对胚胎干细胞特征维持具有重要意义。通过Gather和TFM-Explore两个软件预测中心蛋白编码基因-1200到+200启动子序列的转录因子结合位点,我们发现一个新的转录因子NF-Y能调控9个中心蛋白编码基因。结合已报道的SOX2,OCT-4,NANOG,c-Myc重要转录因子调控靶基因集和I2d蛋白相互作用数据库,我们构建了胚胎干细胞重要转录因子与中心蛋白关系网络图,我们提出一个新的假说:SOX2,OCT-4,NANOG等重要转录因子通过自身调控和相互间调控的回路,从而维持了在胚胎干细胞中合适表达水平,它们进一步通过调控许多重要的中心蛋白编码基因,从而维持了胚胎干细胞相互作用网络拓扑结构稳定,发挥对胚胎干细胞的调节作用。 第三部分:构建人类胚胎干细胞富集的功能相互作用网络。为了鉴定人类胚胎干细胞中富集的功能相互作用方式,我们利用QuickGO软件对人类I2d数据库中13560蛋白进行了基因本体分子功能GOSlim注解,结果表明在13560个蛋白中,目前具有一个及以上的GOSlim分子功能注解蛋白为9881个。进一步采用我们编写的perl程序分别将人类胚胎干细胞富集蛋白相互作用对和I2d蛋白相互作用对注解成相对应的分子功能GOSlim-GOSlim相互作用对,结果:在I2d数据库中的92545对蛋白相互作用对中,共有分子功能GOSlim注解的蛋白相互作用对为74921对;在胚胎干细胞富集蛋白相互作用对中,共有分子功能GOSlim注解的蛋白相互作用对1682对。通过EASE方法分析42个GOSlim组合的903对GOSlim-GOSlim作用对在胚胎干细胞中富集得分,我们发现有66对GOSlim-GOSlim组合在胚胎干细胞中富集。进一步采用Cytoscape软件作图,我们发现除了 GO:0030234之外,其它GOSlim术语均形成一个连续的相互作用网络,在此GOSlim功能相互作用网络中,前4个最高连接的GOSlim术语分别为:GO:0003677,DNA binding;GO:0016787,hydrolase activity; GO:0003723,RNA binding;GO:0003824,catalytic activity。而且,我们研究发现大多数功能相互作用对,涉及转录和翻译过程,这和干细胞自我更新和多潜能维持、无限增殖特性密切相关。 总之,我们的研究已经鉴定了由403个胚胎干细胞高表达基因组成的一个富集蛋白相互作用网络,并进一步发现了一个富集的功能相互作用网络。这些相互作用网络对维持胚胎干细胞功能特征可能具有非常重要的作用,在胚胎干细胞富集蛋白相互作用网络中的中心蛋白,如MYC,H2AFX,RUVBL1,DDX18, CDC2,HDAC2,HIST1H4C等,可能在胚胎干细胞命运决定中具有非常重要作用,值得我们以后进一步深入探讨。但是由于目前蛋白相互作用网络不完全,以及高通量蛋白质组学技术尚不太完善,一些目前已知的重要基因/蛋白如KLF4等尚未包括在我们的蛋白相互作用网络中。尽管如此,我们采用基于复杂网络理论的生物信息学方法对胚胎干细胞富集基因进行了深入探讨,并得到了一些重要的启发,随着蛋白相互作用数据库和蛋白质组学技术的不断完善和改进,我们可能采用相似的方法重新进行网络分析,我们认为,采用不同的方法和技术将会进一步加深对人类胚胎干细胞的认识和了解,加速它的临床应用。 关键词:胚胎干细胞 蛋白相互作用 复杂网路 无标度网络 生物信息学
Human Embryonic Stem Cells (hESCs) are pluripotent cells isolated from the inner cell mass of the blastocyst. There are some specific characteristics of hESCs: (1) Pluripotency. After the removal of differential inhibition, they can differentiate into any kind of tissue cells. (2) Self-renewal and unrestrained proliferation. In suitable conditions, hESCs can self-renew and proliferate stably for a long time, i.e, "immortalization". So, these cells are potentially invaluable in the field of regenerative medicine, pharmacology and auxology. Efforts on this have identified some important genes in hESCs lines, such as OCT3/4, NANOG, REX1, SOX2 and FOXD3. Furthermore, there are some detailed and comprehensive researches of molecule functions on some of these genes and proteins. Molecular biology, based on reductionism thinking style, infers the basic rules about vital processes and vital phenomena by analyzing some individual genes or proteins. However, because conventional molecular biology usually emphasizes the function of single molecule and neglects the interactions between all these molecules, we have no better understanding of the whole complex systems. It is well known that most of the biomolecules such as proteins do not function in isolation, but rather interact with one another to form molecular networks, and larger protein complexes are, in turn, part of a more extensive biological webs. These networks and biological webs determine the characteristics of the cell, the tissue, and even the whole biosystem. Until now, the vast majority of studies about stem cells have only focused on individual genes/proteins, without considering the possible role of interactions. So our knowledge concerning about the characteristics of hESCs is still limited, which also limit the clinical application of stem cells. With the evolution of high-throughput technologies in the post-genomics era, studies have shifted from characterization of single protein to investigation of the entire interactome. Nowadays, biological knowledge is often represented by networks, such as regulatory and metabolic networks. Construction and analyses of these networks have revealed some interesting characteristics within the framework of interactome. There are an increasing number of studies focusing on the transcriptional networks, which emphasize the roles of transcription factors that can regulate human ES cells. As a result, a number of important transcription factors responsible for self-renewal and pluripotency have been identified. In recent years, the network-based approach has gained popularity and been successfully applied for analysis of protein-protein interaction networks in many species and diseases. For instance, one study using affinity purification of NANOG under native conditions followed by mass spectrometry generate a mouse ES protein interaction network including only 37 proteins. However, the picture is far from completeness. Microarray technology provides us a unique opportunity to examine gene expression patterns in hESCs. However, heterogeneity of hESCs gene expression data could exist across different laboratories or different cell lines, which can be partly circumvented by meta-analysis so as to give a more robust result. A "consensus hESCs gene list" is produced by Assou, but some of the genes on the list also express in a few other tissues. The authors explained that these proteins might not be highly specific to hESCs individually, however, they acted together with other proteins to function specifically in hESCs. Based on these observations, we postulate that there may be an enriched protein interaction sub-network(s) among a collection of those on or off the list to maintain or to modify the properties of hESCs. After the discovery of scare free characteristics in the complex network, we have broken through the constraint in random network and found that there is a basic rule in many complex system networks. Complex network theory studies a series of questions, for instance, the network structural feature and the relationship between structure and function in the real network such as society network using graph theory to get more knowledge in real systems. As a rising subject, complex network theory provides a new way and platform for systematic researches on biology networks. Jeong has firstly published a paper focused on topological structure feature of metabolism network using complex network theory in Nature magazine in 2000. Since then, there is a rapid development on studying biology network with the help of the complex network theory. Complex network theory indicates that there is a consistent structure feature between the cellular molecule interaction network and other complex systems networks such as World Wide Web and Social Web, which predicts that there may be a similar rule in controlling the real systems of complex biology network. In this paper, two questions have been raised: (ⅰ) Do the hESC-enriched genes/proteins interact directly with each other more frequently than expected by chance alone? In other word, does a hESC-enriched protein interaction network exist at all? If the answer is affirmative, what characteristics of topology structure are in the network? (ⅱ) Can any enriched functional interaction patterns be identified to be important for maintaining characteristics of hESCs? To address these questions, we perform network-based bioinformatic analysis on both enriched genes/proteins and ontological terms. This thesis can be divided into three parts: 1.In this part, we constructed a protein-protein interaction network and studied the topological characteristics of this network. We collected a list of genes with known entrez gene ID from the original "consensus hESCs gene list" that was published through meta analysis gene chips. And we converted these gene IDs to 1020 UniProtKB/Swiss-Prot accession numbers by UniProt ID mapping and named them hESC-enriched proteins (hESPs). We downloaded the I2d protein interaction database of human and removed reciprocally redundant protein interaction pairs by our peri script, at last we obtained 92545 unique protein interaction pairs that were constituted by 13560 proteins in UniProtKB/Swiss-Prot. We generated 3904 tissue-selective unique protein-coding genes from the latest Affymetrix annotation for selective affymetix probe IDs in 97 normal human tissue that was reported by Prof. Liang Shuang. When comparing hESC-enriched genes with those tissue-selective ones. 274 genes were found intersected. Protein-protein interaction network was constructed by our Perl script which used a breadth-first search algorithm to search the I2d database. This continuous protein interaction network, named by enriching protein-protein interaction network in hESCs because of its larger size than random networks of 1000 samplings, was made up of 403 over-expressed protein in ES cell. Then, this network was visually rendered by Cytoscape program. According to Barabasi-Albert model for complex scare free network, we analyzed the node degree and degree distribution in the network. At last we found out there was a value of 1.3081 for degree exponent, which proved our network was a scare free network. 2.According to the complex network theory, we emphasized and analyzed the hubs proteins in the network. In the first part, we had illuminated there was a scare free feature in our enriching protein interaction network. So we decided to analyze the robustness and vulnerability of the network. In order to do this, we used my own Perl script, which used the Dijkstra algorithm, to calculate the average shortest path length before or after deleting some nodes. We found that the average shortest path length was 3.679、 3.676±0.006、 3.674±0.016, 3.672±0.016、 3.686±0.052 and 3.688±0.040 after deleting 0、4、8、12、16 and 20 random nodes, respectively. Statistical Analysis indicated there was no obvious change for average shortest path length. But there was an obvious change for the average shortest path length after deleting some hubs nodes. The length, which was recorded as 3.679、3.770±0.055、 3.849±0.065、4.028±0.020、4.208±0.118 and 4.448±0.092 after deleting 0、 4、8、 12、16 and 20 hubs proteins respectively, was longer when deleting more hubs nodes. The above result not only demonstrated that there was good robustness and vulnerability in our network but also indicated the hubs proteins played a very important role in maintaining the network's topological structure. According to top 5% highly connected nodes, we have got 21 hubs proteins: MYC、EIF4A1、DDX18、 H2AFX、KIAA0020、RPL4、PCNA、POLR1B、HSPA8、DKC1、CDC2、EIF4E、 BXDC1、BOP1、RPLP0、EIF3A、RUVBL1、HDAC2、GFPT2、HIST1H4C、 CBS. There were some hubs proteins such as MYC、H2AFX and RUVBL1 which have been reported to be involved in self-renewal and proliferation of stem cells. MGI database also indicated there was an important relationship between the embryonic development and some hubs proteins such as POLR1B、CDC2、HDAC2 and MYC, and embryonic lethality would happen if these proteins mutated. Based on these observations, we presumed that these hubs proteins played very important role for embryonic stem cell and deserved to be studied in the future. Through analysis on the promoter sequences of 21 hubs protein-coding genes by Gather and TFM-Explore softwares, we have found a new transcription factor NF-Y might be concerned with stem cells because of its controling 9 hubs proteins. Combining the target gene lists for SOX2, OCT-4, NANOG, c-Myc and I2d database, we constructed a relational network for important transcription factors and hubs proteins in hESCs, and raised a hypothesis that the directed or self-directed regulatory loops of OCT4, NANOG and SOX2 could be used to maintain a proper and stable expression level through a robust synergism. The above reciprocity would futher sustain the stabilization of network topology through regulating many hubs proteins, which was vitally essential for hESCs. 3.We constructed an enriching functional interaction network in hESCs. Interacting partners were likely to be functionally related. For the purpose of identifying functional enrichment in hESCs, we assigned GOSlim function terms to 13560 human UiprotKB/Swiss-Prot proteins in I2d database whenever possible. 42 GOSlim function terms from QuickGO were all used. We found out that 9881 proteins have been annotated by GOSlim function terms. We annotated each hESC-enriched protein-protein interaction pair and I2d protein-protein interaction pair to molecule function GOSlim-GOSlim interaction. As a result, we obtained 74921 protein-protein interaction pairs annotated by GOSlim in I2d database and 1682 protein-protein interaction pairs annotated in hESCs enriching protein interaction pairs. We calculated the enrichment score for each GOSlim-GOSlim interaction pair originated from all human interacting proteins in I2d database or from hESC-enriched proteins based on EASE and found 66 GOSlim-GOSlim interaction pairs enriched in hESCs. After that, we built a network of the above 66 GOSlim-GOSlim interaction pairs by Cytoscape software. In this network, except for G0:0030234, most of the nodes were found in a continuous functionally interacting network. The top fourth most connected GOSlim terms in this network were G0:0003677, G0:0016787, G0:0003723 and G0:0003824. Interestingly, most of these interacting pairs, particularly those formed with G0:0003677 (DNA binding) and GO:0003723 (RNA binding), were involved in transcription and translation. Which was consistent with the need for self-renewal and unrestrained proliferation. In conclusion, this study has identified an enriched protein interaction network formed by 403 hESC-over-expressed gene-coding proteins. Enriched molecular functional interaction network was also found in hESCs. The existence of these interaction networks beyond randomness suggested that they were important and likely to be responsible for the maintenance of hESCs' functions. The hubs governing the hESC-enriched protein interaction network, such as MYC, H2AFX, RUVBL1, DDX18, CDC2, HDAC2 and HIST1H4C, which would possibly play a critical role in determining the fate of hESCs, deserve more attention in future investigations. It is worth noting that some hESC-associated proteins for self-renewal and pluripotency, for instance, KLF4 and other possible factors, were missing in the hESC-enriched protein interaction network due to the incompleteness of protein interaction datas and the shortage of proteomics technology. Despite this, our findings were a step closer toward the right direction, on the systems level, to gain more significant biological information in the stem cell research. When more data becomes available, it will be possible to refine the network and make it more informative for studying hESCs. It is also hoped that continuous researches on hESCs will speed up its therapeutic applications. KEYWORDS: embryonic stem cell; protein interaction; complex network; scare free network; bioinformatics