8th International Conference on Spoken Language Processing

Jeju Island, Korea
October 4-8, 2004

The ICSI-SRI-UW Metadata Extraction System

Elizabeth Shriberg (1,2), Andreas Stolcke (1,2), Dustin Hillard (3), Mari Ostendorf (3), Barbara Peskin (1), Mary Harper (4), Yang Liu (1,4)

(1) ICSI, USA; (2) SRI, USA
(3) University of Washington, USA
(4) Purdue University, USA

Both human and automatic processing of speech require recognizing more than just the words. We describe a state-of-the-art system for automatic detection of "metadata" (information beyond the words) in both broadcast news and spontaneous telephone conversations, developed as part of the DARPA EARS Rich Transcription program. System tasks include sentence boundary detection, filler word detection, and detection/correction of disfluencies. To achieve best performance, we combine information from different types of language models (based on words, part-of-speech classes, and automatically induced classes) with information from a prosodic classifier. The prosodic classifier employs bagging and ensemble approaches to better estimate posterior probabilities. We use confusion networks to improve robustness to speech recognition errors. Most recently, we have investigated a maximum entropy approach for the sentence boundary detection task, yielding a gain over our standard HMM approach. We report results for these techniques on the official NIST Rich Transcription metadata tasks.

Full Paper

Bibliographic reference.  Shriberg, Elizabeth / Stolcke, Andreas / Hillard, Dustin / Ostendorf, Mari / Peskin, Barbara / Harper, Mary / Liu, Yang (2004): "The ICSI-SRI-UW metadata extraction system", In INTERSPEECH-2004, 577-580.