ISCA Archive Odyssey 2004
ISCA Archive Odyssey 2004

Modeling NERFs for speaker recognition

Sachin Kajarekar, Luciana Ferrer, Kemal Sönmez, Jing Zheng, Elizabeth Shriberg, Andreas Stolcke

We introduce a new type of feature to capture long-range patterns associated with individual speakers or with speaking styles . NERFs, or Nonuniform Extraction Region Features, are defined based on regions of speech that are delimited by various automatically extractable events of interest. There is a wide unexplored space of potentially useful NERFs, but to use them successfully, at least two important challenges must be addressed: (1) methods for coping with inherently missing features, and (2) methods for feature selection from large sets of potentially correlated NERFs. We address the issue of missing features in this paper. We propose three methods for modeling NERFs that cope with missing features. We show that on the 2003 NIST extended-data speaker recognition evaluation task, a NERF system yields an EER of 11.6% alone, and improves the MFCC baseline performance by roughly 15% relative.

Cite as: Kajarekar, S., Ferrer, L., Sönmez, K., Zheng, J., Shriberg, E., Stolcke, A. (2004) Modeling NERFs for speaker recognition. Proc. The Speaker and Language Recognition Workshop (Odyssey 2004), 51-56

  author={Sachin Kajarekar and Luciana Ferrer and Kemal Sönmez and Jing Zheng and Elizabeth Shriberg and Andreas Stolcke},
  title={{Modeling NERFs for speaker recognition}},
  booktitle={Proc. The Speaker and Language Recognition Workshop (Odyssey 2004)},