随着智能设备的普及,人机交互已深入渗透人类生产生活的方方面面。情境感知智能人机交互已日益成为人工智能、机器视觉、数据挖掘领域的研究热点。 由于数字化相机的普遍性及其相关技术的通用性,基于单目可见光视觉的人机交互技术已受到越来越多研究者的关注。本文在广泛阅读与调研国内外相关研究的基础上,针对基于单目可见光视觉的情境感知智能人机交互存在的主要问题与不足,开展了一系列的深入研究,提出了如下创新性方法。 提出了一种基于结构相似性时空分析的情境感知光照均衡方法。利用光照补偿结构图与物体光反射特性,获取光照补偿空间分布,通过计算两帧之间非动态物体的光照变化量估计光照时变情况;在基于时间和空间的光照补偿基础上,结合对数直方图均衡算法,实现对视频的快速光照均衡化。实验结果表明,所提方法能够同时改善视频图像的能见度、对比度、自然性、光照一致性和信息稳定性。 提出了一种基于背景建模的复杂情境人体分割方法。根据人体的运动特征及头部结构特征,利用结构相似分布图与头部检测算法,构造出不含人体的帧图像,用于背景模型更新,并采用多特征融合方法,从复杂情境中分割出人体。实验结果表明,所提方法能够实时、准确和完整地获取人体区域。 提出了一种基于差异更新的三维人体骨架估计方法。将人体分割算法与骨架定位算法融合,提高骨架定位效率的同时减少骨架关键点的误检;利用人体前景与人体骨架的运动一致性,基于差异更新算法,抑制人体骨架定位过程中的抖动现象;利用规范化骨架关键点间相对位置特征,建立骨架深度字典模型,获取3D人体骨架信息。实验结果表明所提方法有效、可行。 提出了一种基于模式融合特征点定位的面部朝向估计方法。通过基于双尺度面部区域检测算法定位人的面部区域;利用特征点区域的相对关系,融合两种不同模式的特征点定位算法,提高面部特征点定位的准确性;利用面部特征点这一稀疏特征,建立面部朝向识别模型,实现复杂情境中人脸的面部朝向估计。相比于基于人脸稠密特征的面部朝向估计方法,所提方法具有鲁棒的面部特征描述能力,可有效提高面部朝向估计效果。 提出了一种基于阶段行为特征的交互主体用户感知方法。基于人体部件骨架关键点,阶段性判别人体的指示交互行为状态,识别用户是否具备交互意图;采用空间最邻近算法,从存在交互行为的用户中定位出交互主体用户。实验结果表明所提方法有效、可行。 提出了一种基于躯干位移的交互主体用户跟踪方法。利用人体躯干位移特征跟踪场景中的交互主体用户,采用基于人体骨架区域的色彩直方图匹配算法和快速正面人脸识别算法,恢复由于骨架丢失而中断的用户跟踪链。相比于基于整体区域的跟踪方法,所提方法能更有效地解决人体跟踪过程中区域混叠造成交互用户身份不明确的问题。 提出了一种基于自适应虚拟空间屏的人机交互方法。通过交互主体用户的面部朝向,确定其在交互界面上的关注区域,并将关注区域自适应地映射到虚拟空间屏,通过手部与虚拟空间屏的虚拟接触,响应交互主体用户所表达的交互意图。通过大量的实验和对比分析,结果表明所提出方法能够高效地实现多人有序交互。 关键词:情境感知;人机交互;单目视觉;3D人体骨架;虚拟空间屏
With the popularity of smart devices, human-computer interaction has penetrated into every aspect of human production and life. Context-aware intelligent human-computer interaction has increasingly become a research hotspot in the field of artificial intelligence, machine vision, and data mining. Due to the universality of digital cameras and its related technologies, human-computer interaction technology based on monocular vision has attracted more and more researchers' attention. Some related references at home and abroad have been read and investigated extensively. Human-computer interaction in complex situations is studied in depth aiming at some drawbacks and limitations for context-aware intelligent human-computer interaction based on monocular vision. Some main contributions of this dissertation are as follows: A context-aware illumination equalization method based on spatiotemporal structural similarity is proposed. The illumination compensation structure map and the object light reflection characteristics are utilized to obtain the illumination compensation spatial distribution. The time variation of illumination is estimated by calculating the illumination variation of the non-dynamic object between the two frames. Spatiotemporal illumination compensation and logarithm histogram equalization are developed to fast illumination equalization for video. The experimental results show that the proposed method can simultaneously improve the visibility, contrast, naturalness, illumination consistency, and information stability. A human segmentation method based on background modeling is proposed. According to the motion characteristics of the human body and the features of the head structure, the frame images containing no human body pixels are constructed by using the structural similarity distribution map and the head detection, which are used for background modeling. The background modeling algorithm and multi-features fusion segmentation algorithm are combined to obtain human body from video. The experimental results show that the proposed method can segment human body from video under uncontrolled scene in real time, accurately and completely. A 3D skeleton estimation method based on difference updating is proposed. The above human segmentation method and the skeleton locating algorithm are fused to improve the efficiency of skeleton locating and reduce the false detection of the key points of the skeleton. According to the motion consistency between body and skeleton, the difference update algorithm is developed to restrain the jitter phenomenon of skeleton positioning. Using the normalized relative position feature of the skeleton, the depth dictionary model is built to obtain the 3D skeleton. The experimental results show that the proposed method is effective and feasible. A face orientation estimation method based on pattern fusion facial key point localization is proposed. The two-scale facial region detection algorithm is developed to locate the facial regions from image. Utilizing the relative relationship of the region of facial key points, two different localization algorithms are fused to improve the accuracy of facial key point localization. A face orientation recognition model is built by using facial key points, which describe sparse facial features. Compared with face orientation estimation methods based on facial dense features, the proposed method has the robust ability of facial features description, which can effectively improve the accuracy of face orientation estimation. An interaction subject user perception method based on phased behavior characteristics is proposed. Skeleton points of the parts of human body are used to phased estimate whether the user has the interaction intention. The spatial nearest neighbor algorithm is utilized to locate the interaction subject user from the users who have interactive behavior. The experimental results show that the proposed method is effective and feasible. An interaction subject user tracking method based on human torso displacement is proposed. The feature of human torso displacement is produced to track the interaction subject user. Skeleton regions-based color histogram matching algorithm and fast positive face recognition algorithm are developed to recover the user tracking chain that is interrupted due to skeleton loss. Compared with the tracking methods based on the whole region of human, the proposed method can effectively solve the problem of tracking region overlap, which may cause the identity of user to be confused. A human-computer interaction method based on adaptive virtual space screen is proposed. According the face orientation of the interaction subject user, the interested region on the interaction interface is located. A virtual space screen is produced to map the interested region. Through virtual touch between hand and virtual space screen, the interaction intention of the interaction subject user is responded. Keywords: Context-aware; human-computer interaction; monocular vision; 3D skeleton; virtual space screen