ISCA Archive Eurospeech 1999
ISCA Archive Eurospeech 1999

Adaptation to environment and speaker using maximum likelihood neural networks

Zong Suk Yuk, James Flanagan, Mahesh Krishnamoorthy, Krishna Dayanidhi

When there is a mismatch between training and testing conditions, statistical speech recognition algorithms suffer from severe degra-dation in recognition accuracy. The mismatch could be due to the interference from acoustical environments where systems are actually used or from speakers themselves. In this paper, a neural network based transformation approach is studied to handle the data distribution mismatches between training and testing conditions. The conditional probability that comes from hidden Markov model (HMM) based recognizers is used for the objective function of a neural network. It maximizes the likelihood of the data from a test-ing environment, and allows global optimization of the network when used with HMM-based recognizers. The new objective function can be used to transform speech feature vectors, or the mean vectors and covariance matrices of a recognizer. The proposed algorithm is evaluated on a noisy distant-talking version of the Resource Management database.


doi: 10.21437/Eurospeech.1999-554

Cite as: Yuk, Z.S., Flanagan, J., Krishnamoorthy, M., Dayanidhi, K. (1999) Adaptation to environment and speaker using maximum likelihood neural networks. Proc. 6th European Conference on Speech Communication and Technology (Eurospeech 1999), 2531-2534, doi: 10.21437/Eurospeech.1999-554

@inproceedings{yuk99_eurospeech,
  author={Zong Suk Yuk and James Flanagan and Mahesh Krishnamoorthy and Krishna Dayanidhi},
  title={{Adaptation to environment and speaker using maximum likelihood neural networks}},
  year=1999,
  booktitle={Proc. 6th European Conference on Speech Communication and Technology (Eurospeech 1999)},
  pages={2531--2534},
  doi={10.21437/Eurospeech.1999-554}
}