7th International Conference on Spoken Language Processing

September 16-20, 2002
Denver, Colorado, USA

A Hybrid Approach to Compounds in LVCSR

Tom Laureys (1), Vincent Vandeghinste (2), Jacques Duchateau (1)

(1) K.U. Leuven ESAT-PSI, Belgium; (2) K. U. Leuven CCL, Belgium

In several languages compound words form orthographic units, which complicates the task of ensuring good lexical coverage for large vocabulary continuous speech recognition (LVCSR). A common approach to the problem consists of first recognizing the compound constituents, followed by an automatic recompounding process. We describe an accurate compound module, which combines a rule-based approach with statistical pruning. The module is incorporated in a broadcast news recognition task for Dutch and yields an 11% relative decrease in word error rate (WER).

Full Paper

Bibliographic reference.  Laureys, Tom / Vandeghinste, Vincent / Duchateau, Jacques (2002): "A hybrid approach to compounds in LVCSR", In ICSLP-2002, 697-700.