Cross-Lingual Speaker Adaptation for Statistical Speech Synthesis Using Limited Data

Seyyed Saeed Sarfjoo, Cenk Demiroglu


Cross-lingual speaker adaptation with limited adaptation data has many applications such as use in speech-to-speech translation systems. Here, we focus on cross-lingual adaptation for statistical speech synthesis (SSS) systems using limited adaptation data. To that end, we propose two techniques exploiting a bilingual Turkish-English speech database that we collected. In one approach, speaker-specific state-mapping is proposed for cross-lingual adaptation which performed significantly better than the baseline state-mapping algorithm in adapting the excitation parameter both in objective and subjective tests. In the second approach, eigenvoice adaptation is done in the input language which is then used to estimate the eigenvoice weights in the output language using weighted linear regression. The second approach performed significantly better than the baseline system in adapting the spectral envelope parameters both in objective and subjective tests.


DOI: 10.21437/Interspeech.2016-345

Cite as

Sarfjoo, S.S., Demiroglu, C. (2016) Cross-Lingual Speaker Adaptation for Statistical Speech Synthesis Using Limited Data. Proc. Interspeech 2016, 317-321.

Bibtex
@inproceedings{Sarfjoo+2016,
author={Seyyed Saeed Sarfjoo and Cenk Demiroglu},
title={Cross-Lingual Speaker Adaptation for Statistical Speech Synthesis Using Limited Data},
year=2016,
booktitle={Interspeech 2016},
doi={10.21437/Interspeech.2016-345},
url={http://dx.doi.org/10.21437/Interspeech.2016-345},
pages={317--321}
}