Visualization and Interpretation of Latent Spaces for Controlling Expressive Speech Synthesis Through Audio Analysis

Noé Tits, Fengna Wang, Kevin El Haddad, Vincent Pagel, Thierry Dutoit


The field of Text-to-Speech has experienced huge improvements last years benefiting from deep learning techniques. Producing realistic speech becomes possible now. As a consequence, the research on the control of the expressiveness, allowing to generate speech in different styles or manners, has attracted increasing attention lately. Systems able to control style have been developed and show impressive results. However the control parameters often consist of latent variables and remain complex to interpret.

In this paper, we analyze and compare different latent spaces and obtain an interpretation of their influence on expressive speech. This will enable the possibility to build controllable speech synthesis systems with an understandable behaviour.


 DOI: 10.21437/Interspeech.2019-1426

Cite as: Tits, N., Wang, F., Haddad, K.E., Pagel, V., Dutoit, T. (2019) Visualization and Interpretation of Latent Spaces for Controlling Expressive Speech Synthesis Through Audio Analysis. Proc. Interspeech 2019, 4475-4479, DOI: 10.21437/Interspeech.2019-1426.


@inproceedings{Tits2019,
  author={Noé Tits and Fengna Wang and Kevin El Haddad and Vincent Pagel and Thierry Dutoit},
  title={{Visualization and Interpretation of Latent Spaces for Controlling Expressive Speech Synthesis Through Audio Analysis}},
  year=2019,
  booktitle={Proc. Interspeech 2019},
  pages={4475--4479},
  doi={10.21437/Interspeech.2019-1426},
  url={http://dx.doi.org/10.21437/Interspeech.2019-1426}
}