深度学习技术极大地促进了语音领域的科学研究进展,推动了语音输入法、车载语音助手和家庭智能语音设备等产品的市场化,自然人机语音交互迎来了广泛的应用需求。在自然人机语音交互场景中,人与机器之间的交互一般不再受到距离的限制。当人与机器之间距离较远时,环境中的背景噪声、空间混响和潜在的语音信号重叠会显著降低目标语音的质量,进而影响语音通信和识别交互系统的鲁棒性。如何抑制噪声、混响和非目标语音的消极作用,并提高实际复杂场景下远场拾音的质量是本文的研究重点。 在远场拾音方法中,麦克风阵列技术由于采用多个观测通道,在信号时频信息之外还可以感知声源的空间信息,因而具有更高的研究和实用价值。经典的阵列信号处理方法众多,包括声源定位、线性滤波波束形成、混响抑制和盲源分离等方法,然而这些方法一般对声学场景有着较强的假设,如假设噪声相对于语音信号更加平稳,或假设各个信号源之间统计独立。当这些假设不成立时,以上算法的性能会受到显著影响。因此,非平稳噪声和多声源等复杂声学场景下的语音拾取还面临诸多问题和挑战。与此同时,深度学习技术通过收集先验数据,训练具有特定结构的深度神经网络模型,可以引入关于噪声特点、声源频谱和声源位置的先验信息。将这些先验信息与阵列信号处理算法相结合,形成了新的结合了深度学习与麦克风阵列的远场拾音解决方案。 本文即通过分析远场拾音场景下的目标声源定位和目标声源拾取等问题,采用以阵列信号处理为核心、结合深度学习技术的研究方案,探索了新的声源定位与线性滤波方法,提高了实际复杂场景下的远场拾音质量以及后端语音识别系统性能。主要研究工作及创新点包括: 1.研究了多声源场景下的目标声源定位问题,提出了基于混合沃森模型和时频选择网络的目标声源定位方法。根据信号频谱的稀疏性假设,在目标声源受到噪声或其他语音干扰时仍然存在目标声源占优的时频点,准确判断出这些时频点即可以实现精确定位。特别地当环境中存在多个声源时,以往的定位方法无法直接确定目标声源的方位,一般需要额外的辅助信息来进行后续处理。本文设计了目标时频掩蔽估计深度神经网络,提出利用目标声源的一句话作为先验信息,引导掩蔽估计网络只关注于目标声源,从而选择出目标声源占优的时频点用于定位判断。同时在定位算法中引入混合沃森模型来建模观测信号,同时利用了阵列通道间的相位差与能量差,在存在非目标语音干扰和低信噪比的场景中有效地提高了定位结果的准确度。 2.研究了日常生活场景下基于多通道线性滤波的目标声源拾取问题。首先,本文从理论上分析了多通道维纳滤波与可扩展子空间滤波、最大化信噪比滤波之间的区别与联系,在此基础上提出了具有频带平稳残余噪声输出特点的多通道维纳滤波器。其次,结合基于深度神经网络的时频掩蔽估计方法,实现了语音和噪声协方差矩阵的鲁棒估计和滤波器系数的计算。本文分析发现,由于混响和实际掩蔽估计误差的影响,频域窄带信号模型假设下的语音协方差矩阵秩为1的结论并不成立。为此提出了基于广义特征分解的语音信号协方差矩阵秩1重构方法,提高了滤波器在单目标声源情形下的噪声抑制能力,并推广适用于所有线性滤波器。在包含咖啡馆、公交车和步行街等日常生活场景的实录数据上,提出的方法极大地降低了后端语音识别系统的词错误率。 3.提出了基于声源位置与深度神经网络的相对传递函数建模与预测方法。相对传递函数描述了一对麦克风对于声源的冲激响应之间的相互关系,在上述的声源定位和线性滤波算法中都发挥着重要作用。相对传递函数一般根据观测信号进行估计,而以往先验信息的获取只能依赖自由场假设下的直达声模型。本文利用深度神经网络以数据驱动的方式建模,并基于声源位置预测相对传递函数的先验信息。在首先试验了全监督训练网络的基础上,通过分析稳定声学场景下相对传递函数的低维包络与局部线性特点,建立了只依赖少量标注数据的半监督学习模型。将先验的相对传递函数信息用于广义旁瓣抵消算法,展现出了一定的实用价值。 关键词:远场拾音,麦克风阵列,深度学习,线性滤波,声源定位
The deep learning technology has greatly promoted the recent developments in the speech research field, and it has accelerated the popularization of voice input, voice assistant in car and voice advices in smart home. Now there is an increasing demand for natural human machine interaction. In such cases, one ideally can interact with machine at any distance. But when the distance becomes longer, the speech is unavoidably corrupted by background noise, room reverberation and potential speech overlap. Thus the speech quality is degraded and the succeeding speech communication and interaction system is affected. The main focus of this paper is therefore on suppressing the background noise, room reverberation and non-target speech to achieve better far-field speech processing speech quality in adverse environments. In far-field speech processing, the microphone array technology draws more attention, since it takes multiple observations of the source and makes use of both the spectral and the spatial properties of the signal. The classical methods include source localization, linear filtering/beamforming, dereverberation, blind source separation and so on. Generally, these methods are based on certain assumptions of the acoustic environment, such as that noise is more stationary than speech, and that the signals in the mixture are mutually independent. Once the assumptions are not fulfilled, their performance drops largely in practice. There is still a long way to go for high quality speech processing in scenarios containing non-stationary noise and multi-talker. Meanwhile, the deep learning technology trains carefully designed deep neural networks based on pre-collected data, and hence establishes prior knowledge of the noise, the speech spectrum and the source position. Integrating the prior knowledge with signal processing methods, leads to new far-field speech processing solutions based on deep learning and microphone array techniques. After analysing the issues encountered in far-field target source localization and speech capturing, this dissertation integrates deep learning with microphone array farfield speech processing, and proposes novel source localization and linear filtering algorithms, to improve the speech quality and speech recognition performance in real environments. The main contributions are listed as follows. 1.Research on target source localization in multi-talker scenarios. A target source localization algorithm based on complex Watson mixture model and time-frequency selection neural network is proposed. Based on the sparsity assumption of the signal spectrum, there exist target-dominant time-frequency bins that can be used for localization, even when the target is corrupted by noises and interferences. The challenge is hence to distinguish these bins. Traditional method could not achieve direct target localization without the help of additional information of the target. Notably, an utterance from the target is employed here to guide a mask estimation neural network. The neural network then adapts to the target and output only target-dominant masks, which are subsequently used with a complex Watson mixture model for localizing the target. This model utilizes both the interchannel phase difference and interchannel level difference. Experiments in scenarios with competing interference and high noise levels confirm the effectiveness of the proposed method. 2.Research on multichannel linear filtering based speech processing in daily environments. Firstly, the equivalent conditions for multichannel Wiener filter and variable span filter, maximum signal-to-noise ratio filter are derived in theory. A novel filter that outputs constant residual noise power is then proposed. Secondly, mask estimation neural networks are employed for calculation of the speech/noise spatial covariance matrixes and the filter coefficients. It is found that due to reverberation and estimation error, the speech covariance matrix is not of rank 1 indeed, which violates the frequency-domain signal model under narrowband assumption. A rank-1 reconstruction of the speech covariance matrix using a generalized eigenvector is then proposed. It enhances the noise suppression ability of all the known filters in the single target cases. The proposed method largely reduces the speech recognition errors in daily environments, such as cafeteria, bus and pedestrian area. 3.Research on relative transfer function modelling and prediction based on source pose. Relative transfer function relates the two impulse responses of a pair of microphones to one source. It is an essential variable in the above source localization and linear filtering methods. The estimation of relative transfer function is usually based on the observed signals, while without making an observation, the prior knowledge of relative transfer function comes only from the direct sound model in free field. A pure data-driven method relying on deep neural network is proposed to predict the relative transfer function given the source pose. After experiments on a fully-supervised learning setup, a semi-supervised model is designed based on the low-dimension manifold property of the relative transfer function. The model requires few labelled training data. Applying the predicted relative transfer function in the generalized sidelobe canceller, validates the usefulness of the proposed method. Keywords: Far-field Speech Processing, Microphone Array, Deep Learning, Linear Filtering, Source Localization