空间内涵是信息的重要构成要素,研究表明,在人类社会中有超过57%的数据信息与地理位置、空间区位相关。特别是随着移动定位、传感器网络、移动互联等技术的发展,在人们日常生活生产实践中,通过网络环境传输与共享的泛在地理信息呈指数级增长,其构成了对客观世界多尺度、大纵深、全覆盖的动态映像。在互联网泛在地理信息中,网络文本是其重要的存在形式,有近20%的网络文本中包含对地理位置信息的描述,并且超过四分之一的网络文本检索与地理位置相关。网络文本中地理位置描述的大量存在以及人们对地理位置信息的普遍需求,使得空间内涵成为了信息解析与表达过程中需要考虑的核心内容。 面对包含丰富地理位置描述且数量庞大、结构复杂的网络文本,如何从地理空间视角人手,实现对网络文本信息的挖掘认知是当前地理信息科学领域面临的重大挑战。一方面,网络文本主要基于非结构化的自然语言进行描述,如何从中有效提取出结构化的空间以及相应时间、语义维度的信息内容,并形成对客观世界事件、过程、现象等的集成整合表达,是网络文本理解认知过程中需要探索的重要内容;另一方面,网络文本中的地理位置内涵为信息的空间认知提供了重要视角,如何有效利用该视角,以地图可视化的形象思维模式,将挖掘提取出的位置关联信息内容基于地图进行传输与表达,也是亟待探索的重要课题。 在上述网络文本的信息提取与可视化表达的实际问题驱动之下,本研究着重从信息的空间、时间、语义三个基本维度入手,对网络文本信息进行结构化提取与建模,并将建模后的信息进行集成整合;在此基础上,以地图作为信息传输与表达的载体,结合地图制图学、信息可视化、地图混搭等理论与方法,实现集成整合后的信息向地图空间的映射与可视化表达。本研究中主要从以下几个方面进行了探索: (1)针对包含地理位置内涵的网络文本数据,提出了基于空间、时间、语义基本维度的位置关联信息模型框架,结合地理对象和地理事件两个视角,对自然语言描述的网络文本信息进行形式化表达。在该模型框架的基础之上,利用自然语言处理、命名实体识别、地理编码等技术手段,进行网络文本的时空语义信息结构化提取与解析。并对解析结果中的地理位置描述进行了进一步的地理位置歧义消除、地理位置焦点获取等研究工作,从而完成网络文本的位置关联信息建模。位置关联信息模型的构建为实现信息向地图空间的映射与可视化表达提供了支撑。 (2)研究和探索了以地图作为信息传输与表达的载体,利用地图混搭的模式,实现信息基于地图的索引浏览与空间认知。在此过程中,为了解决信息向地图空间映射与混搭呈现时的信息过载以及图面视觉混乱的问题,研究中对信息项之间的时空语义相似度进行度量,完成信息内容的聚类与集成整合;同时对地图学、信息可视化中的理论与方法进行了探索与拓展,将图内标注与边界标注两种信息图面呈现模式应用于地图混搭框架中,实现信息的图面优化配置与动态交互式呈现。研究通过对位置关联信息的内容组织与呈现方式进行优化,为信息的地图可视化提供了有效模式。 (3)除了上述以单个信息项作为基本单元,对信息基于空间维度的地图浏览与认知进行探索之外,研究中还对隐含于信息集群中实体层级的知识进行了挖掘与地图可视化表达。研究具体以信息集群中所涉及的实体为分析对象,将实体在信息中的共现关系作为依据进行实体关系网络图构建,然后基于该关系图进行实体的重要性度量,并结合网络图边绑定算法对实体关系进行可视化。研究基于上述内容,实现了信息集合中实体关系的挖掘与直观呈现,为信息的实体层级知识的可视化表达提供了有效策略。 关键词:网络文本。位置关联信息,地图可视化,地图混搭,文本挖掘
Space is an important organizational unit of information. Research shows that nearly 57% of information in human society is related to spatial location Especially with the development of GPS, sensor network, mobile internet etc, the ubiquitous geographic information transmitted and shared through the network grows exponentially, which constitutes a dynamic image of the objective world with multiple scales, large depth and full coverage. In the ubiquitous geographic information on the Internet, web text is the most important form of its existence. Nearly 20% of the web text contains the geographical location information, and more than a quarter of the network retrieval is related to geographical location The large number of geographical location description in web text and people's general demand for geographic location information make spatial content become the core factor in the process of information extraction and analysis. In the face of the complex network texts containing rich geographical location descriptions, how to perceive and resolve web text information from the perspective of space is a major challenge facing the field of geographic information science. On the one hand, the web text is mainly formed in unstructured natural language, how to effectively extract structured information content from the spatial, temporal and semantic perspectives is an important issue which needs to be explored; On the other hand, the geographical location content provides an important perspective for the spatial cognition of web textual information How to make effective use of this perspective to transmit and express the extracted location-referenced information based on the map is also an important topic to be explored. Driven by the practical problems of information extraction and visual expression of the web text, this study focuses on the spatial, temporal and semantic dimensions to carry out structured extraction and formal modeling of web textual information On this basis, the map as a carrier of spatial information, is combined with cartographic processing and information visualization methods, to achieve the visual expression of location-referenced web textual information. The research is carried out from the following aspects: (1)A model of location-referenced web textual information based on the spatial, temporal and semantic dimensions is proposed, from which the web text information formed in natural language is formally expressed. On the basis of the model, the web text is extracted and analyzed by natural language processing, named entity recognition, geocoding. Furthermore, the location descriptions in the model are further analyzed, such as the elimination of geographic ambiguity and the acquisition of geographic focus, so as to complete the modeling of location-referenced web textual information. The construction of information model provides support for the visualization of information in the map space. (2)The specific forms and methods for map-based information indexing and browsing are explored, by taking map mashup as the form of location-referenced information presentation. In this process, in order to solve the problem of information overload and visual confusion in the map mashup results, the spatial, temporal and semantic similarity between information items was measured to cluster and integrate the information set. At the same time, the theories and methods in cartography and information visualization are explored and expanded, and two kinds of information presentation modes, namely, traditional labeling and boundary labeling, are applied to optimize the map mashup effect By exploring the content organization and presentation mode of location-referenced information, an effective framework is provided for map-based information browsing and spatial cognition (3)The entity level knowledge contained in the location-referenced information set is explored in this research. Specifically, the entities involved in the information are taken as the analysis objects, and co-occurrences of entities in the information are analyzed to construct the entity-relationship graph. Based on the constructed entity-relationship graph, the importance of entities are measured and edge bundling algorithm is used to visually present the entity-relationship. Thus, an effective strategy for the entity level knowledge discovery and expression for the information set is provided in this research. Keywords: Web textual information, Location-referenced information, Cartographic visualization, Map mashups, Text mining