ISCA Archive Interspeech 2009
ISCA Archive Interspeech 2009

Probabilistic and possibilistic language models based on the world wide web

Stanislas Oger, Vladimir Popescu, Georges Linarès

Usually, language models are built either from a closed corpus, or by using World Wide Web retrieved documents, which are considered as a closed corpus themselves. In this paper we propose several other ways, more adapted to the nature of the Web, of using this resource for language modeling. We first start by improving an approach consisting in estimating n-gram probabilities from Web search engine statistics. Then, we propose a new way of considering the information extracted from the Web in a probabilistic framework. Then, we also propose to rely on Possibility Theory for effectively using this kind of information. We compare these two approaches on two automatic speech recognition tasks: (i) transcribing broadcast news data, and (ii) transcribing domain-specific data, concerning surgical operation film comments. We show that the two approaches are effective in different situations.


doi: 10.21437/Interspeech.2009-128

Cite as: Oger, S., Popescu, V., Linarès, G. (2009) Probabilistic and possibilistic language models based on the world wide web. Proc. Interspeech 2009, 2699-2702, doi: 10.21437/Interspeech.2009-128

@inproceedings{oger09_interspeech,
  author={Stanislas Oger and Vladimir Popescu and Georges Linarès},
  title={{Probabilistic and possibilistic language models based on the world wide web}},
  year=2009,
  booktitle={Proc. Interspeech 2009},
  pages={2699--2702},
  doi={10.21437/Interspeech.2009-128}
}