当前,随着网络技术的飞速发展和网络用户的海量增加,在线社会网络和社交媒体已经在信息扩散和传播过程中扮演着至关重要的角色。在线社会网络中的信息扩散有着诸多与传统媒介显著不同的特点,吸引了大量研究人员对此开展研究。这些研究能够帮助商业人士制定更加有效的市场营销决策,帮助用户更容易发现有价值的内容,帮助政府有关部门及时发现和抑制有害内容的传播。 虽然已有研究已经取得了众多的研究成果,但是,对于影响社会网络中信息扩散的一些重要因素以及信息扩散模型的研究仍然存在着诸多的不足。为此,本文从寻找和定量分析对于信息传播冇重要作用的影响因素,以及社会网络中的信息扩散建模这两个方面出发开展工作,论文的主要研究内容和成果包括: 在影响社会网络信息扩散的结构因素方面,本文提出了一个“葡萄”结构模型用于描绘社会网络中大规模信息扩散的总体发生过程,并提出了信息扩散的四个结构模式,每个模式都代表一种典型的级联结构并且反映出不同的扩散机制。本文收集了45000多个大规模社会网络信息扩散事件,它们的扩散规模从几千到几十万个节点不等。通过对这些扩散事件的分析,本文发现大多数的大规模扩散的网络结构是非常稀疏且缺乏聚集属性的,其扩散过程中具有两个令人意外的共性现象:首先,即便是这些传播范围很广的热门事件,它们在人群中的“感染概率”仍然非常的低;其次,网络中一个节点被某个事件“感染”的概率并不会因为这个事件被反复曝光而增加。这些现象有效支持了本文提出的“葡萄”结构模型。最后,本文将大规模扩散与小规模扩散做了对比,发现结构特性并不是区分大规模扩散和小规模扩散的关键因素。 在影响大规模社会网络中的信息扩散的内容因素方面,本文首先根据内容本身以及用户行为数据提取了多项特征用于表示内容自身的吸引力。在用户转发行为预测任务上的实验表明,利用内容吸引力特征能够显著提升实验的预测效果,说明其对于社会网络信息扩散具有重要影响。其次,本文采用多种方式构建主题模型推导传播内容的主题分布以及用户的兴趣偏好,并计算了用户兴趣与内容之间的匹配度。在用户转发行为预测实验中,本文分别采用不同方式计算得到的内容兴趣匹配度作为特征,能够获得7%到14%的预测效果提升。这些实验结果表明内容兴趣匹配度对信息扩散具有显著影响。 在影响大规模社会网络中的信息扩散的外部影响力方面,本文设计了一个重构级联树算法,能够将信息传播过程中所受外部影响力的作用抽离出来。通过将算法应用于海量扩散数据并进行统计分析,本文得到关于外部影响力的四个有价值的发现:首先,尽管只有小部分的传播是由外部影响力直接导致的,但在整个扩散过程中外部影响力有着非常重要的作用。具体而言,有50%到70%的扩散节点都受到了外部影响力的作用,并且随着级联扩散规模的增大,这种影响会随之变强。其次,外部影响力能够促使用户和网络中的陌生人发生交互,让更多的用户之间建立连接,社会网络的网络结构因此变得更加紧密,从而有益于信息的扩散。再次,外部影响力可以分为两种,一种影响扩散的规模而另一种偏向于影响扩散的深度。最后,由于外部影响力的存在,传统的意见领袖的作用会下降,普通用户创造的优质内容更有机会被快速传播。 在上述信息扩散影响因素分析的基础上,本文提出了用于建模社会网络信息扩散的扩散-潜在狄利克雷配置(Diffusion-Latent Dirichlet Allocation,D-LDA)模型,它在同一个生成过程中融合了内容的主题分布计算以及用户间影响力计算。模型通过两个步骤交替迭代推导:U-Part推导得到内容与用户的主题分布,而D-Part则计算得到用户之间的影响力以及内容自身的吸引力。D-LDA模型能够区分用户转发行为中所包含的内容相关因素和影响力相关因素。模型训练采用Gibbs Sampling进行参数推断。最后,在大规模数据上通过多项实验对于D-LDA模型的结果进行评测。实验结果表明D-LDA模型能够在少量迭代后迅速收敛,在多个不同评估任务和指标上均取得了比基线模型更好的性能。 关键词:社会网络 信息扩散 内容吸引力 外部影响力 扩散-潜在狄利克雷配置模型
With the rapid development of internet technology and the fast growth of internet users, online social network and social media websites have already become a major role in the spread of information at present. The information dissemination in online social networks has many characteristics that are significantly different from traditional mass media, which has attracted a large number of researchers to cany out related research. The research of information diffusion in online social networks can help business people make more powerful marketing strategies, help users more easily discover valuable information and help government restrict the spread of harmful content. Although the existing research has achieved a lot of results, there are still many deficiencies in the study of information diffusion in social networks, such as identifying key factors in information diffusion and modeling the information diffusion process. The thesis aims to provide insights into information diffusion on online social networks in two ways: to quantitatively analyze the key factors that affect information diffusion, and to model the information diffusion process. Specifically, the main contributions of the thesis are presented as follows. The first part studies the structure and evolution of very large cascades in online social network. We propose the "grapes" model which roughly describe the diffusion process of large cascades and summarizes four diffusion structural patterns, each of which represents a typical cascade structure and reflects different diffusion mechanism. We investigate over 45000 large cascades, and their sizes range from thousands to hundreds of thousands. We found that the underlying network structure of most large cascades are fairly sparse and less clustering, and there are two surprising common phenomena in these diffusion process: first, even if these popular events spread widely, their "infection rate" are still very small; second, the probability of a node being infected does not increases linearly but be somewhat persistence with the number of repeated exposures. These phenomena significantly support the "grapes" model. Finally, we made a comparison between large and small cascades. The results suggest that the structure features are not a key factor in predicting the future growth of cascades. The second part studies the effect of the content of the message on the information propagation. A number of features are extracted based on the text of content and user behavior data to indicate the attractiveness of the content. The experiment on user repost behavior prediction task shows that content attractiveness has a significant influence on the information diffusion in social network. We also propose several schemes to train the topic model to infer user's preference and the matching degree between the user's interest preferences and content. Using these matching degrees as features, it can get a 7% to 14% improvement in user behavior prediction. The experimental results demonstrate that the matching degree plays an important role in the diffusion process. The third part measures the effects of external influence in Information diffusion. We develop an algorithm which allows us to distinguish the effects of external influence in diffusion process. By applying the algorithm to millions of diffusion cascades, we get four valuable findings. First, although only a small portion of reshare activities arise from external influence directly, external influence plays a significant role in information diffusion. In particular, external influence affects nearly 50% to 70% of cascade node in average, and the effects become stronger as the cascade becomes larger. In addition, external influence motivates users to reshare from strangers and improve the odds of being friends between them, which will lead the underlying network denser and benefit the diffusion process. Furthermore, we characterize external influence as two categories: one category mainly affects the size of cascade tree and the other focuses on affecting the depth. At last, we find that, due to the external services, the influentials become less important and more large cascades can be triggered by ordinary people. Finally, based on the above analysis of the key factors of information diffusion on large-scale social networks, we propose a novel Diffusion-Latent Dirichlet Allocation (D-LDA) model, which integrates both content topic inference and social influence computing in the same generative process. The iteration of this model alternates between two steps: U-Part step infers the content and users' topic distribution, and D-Part step calculates the peer's influence and the attractiveness of content. The model has the ability to distinguish the content-related and influence-related factors which affect the user reshares the content. We use Gibbs Sampling to derive the parameters of the model. The D-LDA model is evaluated through multiple experiments on large-scale dataset. The experimental results show that the D-LDA model can converge quickly after a few iterations and produces significantly higher quality results than the prior models. KEY WORDS: social network, information diffusion, attractiveness of the content, external influence, Diffusion-Latent Dirichlet Allocation