9th Annual Conference of the International Speech Communication Association

Brisbane, Australia
September 22-26, 2008

Soft Margin Estimation with Various Separation Levels for LVCSR

Jinyu Li (1), Zhi-Jie Yan (2), Chin-Hui Lee (1), Ren-Hua Wang (2)

(1) Georgia Institute of Technology, USA; (2) University of Science & Technology of China, China

We continue our previous work on soft margin estimation (SME) to large vocabulary continuous speech recognition (LVCSR) in two new aspects. The first is to formulate SME with different unit separation. SME methods focusing on string-, word-, and phone-level separation are defined. The second is to compare SME with all the popular conventional discriminative training (DT) methods, including maximum mutual information estimation (MMIE), minimum classification error (MCE), and minimum word/phone error (MWE/MPE). Tested on the 5k-word Wall Street Journal task, all the SME methods achieves a relative word error rate (WER) reduction from 17% to 25% over our baseline. Among them, phone-level SME obtains the best performance. Its performance is slightly better than that of MPE, and much better than those of other conventional DT methods. With the comprehensive comparison with conventional DT methods, SME demonstrates its success on LVCSR tasks.

Full Paper

Bibliographic reference.  Li, Jinyu / Yan, Zhi-Jie / Lee, Chin-Hui / Wang, Ren-Hua (2008): "Soft margin estimation with various separation levels for LVCSR", In INTERSPEECH-2008, 269-272.