When accented speech beats native one?

Most TTS systems are built to sound as native as possible. Together with Li Tang(Doctoral program), and Prof. Julian Villegas trained a system to sound like a Japanese person speaking English.
We did this by taking an English speech synthesizer and fine-tuning it using recordings from Japanese learners of English. Then, we tested it on 14 native Japanese listeners. The results were encouraging: Listeners understood the Japanese-accented sentences better than the native-accented ones, and they consistently recognized the accent as Japanese.
This work is a step toward real-time accent conversion. Take air traffic control, for example: English is the mandated language for communication. Adapting the accent of an instruction to match what a controller or pilot is used to hearing could ease communication, without the extra time and complexity a translation step would add.

This work received the Best Application Paper Award on September 24, 2026, at the 2026 International Conference on Awareness Science and Technology (iCAST).


s-Weixin Image_20260926153318_98_416.jpg

URL conference: https://icast2026.org.cn/
URL lab: https://onkyo.u-aizu.ac.jp/