Metrics for Modeling Code-Switching Across Corpora

Gualberto Guzmán, Joseph Ricard, Jacqueline Serigos, Barbara E. Bullock, Almeida Jacqueline Toribio


In developing technologies for code-switched speech, it would be desirable to be able to predict how much language mixing might be expected in the signal and the regularity with which it might occur. In this work, we offer various metrics that allow for the classification and visualization of multilingual corpora according to the ratio of languages represented, the probability of switching between them, and the time-course of switching. Applying these metrics to corpora of different languages and genres, we find that they display distinct probabilities and periodicities of switching, information useful for speech processing of mixed-language data.


 DOI: 10.21437/Interspeech.2017-1429

Cite as: Guzmán, G., Ricard, J., Serigos, J., Bullock, B.E., Toribio, A.J. (2017) Metrics for Modeling Code-Switching Across Corpora. Proc. Interspeech 2017, 67-71, DOI: 10.21437/Interspeech.2017-1429.


@inproceedings{Guzmán2017,
  author={Gualberto Guzmán and Joseph Ricard and Jacqueline Serigos and Barbara E. Bullock and Almeida Jacqueline Toribio},
  title={Metrics for Modeling Code-Switching Across Corpora},
  year=2017,
  booktitle={Proc. Interspeech 2017},
  pages={67--71},
  doi={10.21437/Interspeech.2017-1429},
  url={http://dx.doi.org/10.21437/Interspeech.2017-1429}
}