15th Annual Conference of the International Speech Communication Association

September 14-18, 2014

Application of Convolutional Neural Networks to Speaker Recognition in Noisy Conditions

Mitchell McLaren, Yun Lei, Nicolas Scheffer, Luciana Ferrer

SRI International, USA

This paper applies a convolutional neural network (CNN) trained for automatic speech recognition (ASR) to the task of speaker identification (SID). In the CNN/i-vector front end, the sufficient statistics are collected based on the outputs of the CNN as opposed to the traditional universal background model (UBM). Evaluated on heavily degraded speech data, the CNN/i-vector front end provides performance comparable to the UBM/i-vector baseline. The combination of these approaches, however, is shown to provide improvements of 26% in miss rate to considerably outperform the fusion of two different features in the traditional UBM/i-vectors approach. An analysis of the language- and channel-dependency of the CNN/i-vector approach is also provided to highlight future research directions.

Full Paper

Bibliographic reference.  McLaren, Mitchell / Lei, Yun / Scheffer, Nicolas / Ferrer, Luciana (2014): "Application of convolutional neural networks to speaker recognition in noisy conditions", In INTERSPEECH-2014, 686-690.