8th International Conference on Spoken Language Processing

Jeju Island, Korea
October 4-8, 2004

Modeling and Recognition of Phonetic and Prosodic Factors for Improvements to Acoustic Speech Recognition Models

Sarah Borys, Aaron Cohen, Mark Hasegawa-Johnson, Jennifer Cole

University of Illinois, USA; University of Illinois, USA

This paper examines the usefulness of including prosodic and phonetic context information in the phoneme model of a speech recognizer. This is done by creating a series of prosodic and phonetic models and then comparing the mutual information between the observations and each possible context variable. Prosodic variables show improvement less often than phone context variables, however, prosodic variables generally show a larger increase in mutual information. A recognizer with allophones defined using the maximum mutual information prosodic and phonetic variables outperforms a recognizer with allophones defined exclusively using phonetic variables.

Full Paper

Bibliographic reference.  Borys, Sarah / Cohen, Aaron / Hasegawa-Johnson, Mark / Cole, Jennifer (2004): "Modeling and recognition of phonetic and prosodic factors for improvements to acoustic speech recognition models", In INTERSPEECH-2004, 3013-3016.