ISCA Archive SLAM 2013
ISCA Archive SLAM 2013

LMELECTURES: a multimedia corpus of academic spoken English

Korbinian Riedhammer, Martin Gropp, Tobias Bocklet, Florian Hönig, Elmar Nöth, Stefan Steidl

This paper describes the acquisition, transcription and annotation of a multi-media corpus of academic spoken English, the LMELectures. It consists of two lecture series that were read in the summer term 2009 at the computer science department of the University of Erlangen- Nuremberg, covering topics in pattern analysis, machine learning and interventional medical image processing. In total, about 40 hours of high-definition audio and video of a single speaker was acquired in a constant recording environment. In addition to the recordings, the presentation slides are available in machine readable (PDF) format. The manual annotations include a suggested segmentation into speech turns and a complete manual transcription that was done using BLITZSCRIBE2, a new tool for the rapid transcription. For one lecture series, the lecturer assigned key words to each recordings; one recording of that series was further annotated with a list of ranked key phrases by five human annotators each. The corpus is available for non-commercial purpose upon request.

Index Terms: corpus description, academic spoken English, e-learning


Cite as: Riedhammer, K., Gropp, M., Bocklet, T., Hönig, F., Nöth, E., Steidl, S. (2013) LMELECTURES: a multimedia corpus of academic spoken English. Proc. First Workshop on Speech, Language and Audio in Multimedia (SLAM 2013), 102-107

@inproceedings{riedhammer13_slam,
  author={Korbinian Riedhammer and Martin Gropp and Tobias Bocklet and Florian Hönig and Elmar Nöth and Stefan Steidl},
  title={{LMELECTURES: a multimedia corpus of academic spoken English}},
  year=2013,
  booktitle={Proc. First Workshop on Speech, Language and Audio in Multimedia (SLAM 2013)},
  pages={102--107}
}