EventsThe 5th International Electronic Conference on Applied Sciences
Published
This submission belongs to the session S3. Computing and Artificial Intelligence of the event The 5th International Electronic Conference on Applied Sciences
Published date
02 Dec, 2024
Academic Editor
author-avatarEugenio Vocaturo
Citation
Raja Hashim Ali, Muhammad Ramiz Saud, Muhammad Romail Imran, Towards a More Natural Urdu: A Comprehensive Approach to Text-to-Speech and Voice Cloning, in Proceedings of The 5th International Electronic Conference on Applied Sciences, 4 December–6 December 2024, MDPI: Basel, Switzerland
Share
Email
Facebook
Twitter
LinkedIn

Towards a More Natural Urdu: A Comprehensive Approach to Text-to-Speech and Voice Cloning

1. Department of Artificial Intelligence, Faculty of Computer Science and Engineering, Ghulam Ishaq Khan Institute of Engineering Sciences and Technology, 23460 Topi, Khyber Pakhtoonkha, Pakistan., Pakistan
2. Department of Business, University of Europe for Applied Sciences, Think Campus, 14469 Potsdam, Germany., Germany
3. Artificial Intelligence Research (AIR) Group, , Department of Artificial Intelligence, Faculty of Computer Science and Engineering, Ghulam Ishaq Khan Institute of Engineering Sciences and Technology, 23460 Topi, Khyber Pakhtoonkha, Pakistan.
Abstract

This work focuses on the centrality of NLP and TTS in promoting communication for the Urdu-speaking population where there is a dearth of language assets in the regional languages. While English and other languages of European origin have reliable computational assets available, Urdu is still considered relatively illiterate in this aspect, and hence restricted.
Therefore, to address this problem, we constructed our own dataset using audio from a YouTube playlist that contains an Urdu novel reader for more than 100 hours. This dataset was carefully preprocessed for our use and different errors were corrected to provide high-quality input for our TTS models. Our work constitutes one of the first research attempts at creating a large-scale Urdu speech dataset and at employing unique techniques of Automatic Speech Analysis. To achieve this purpose, the linguistic and cultural characteristics of the Urdu language are incorporated in this approach to guarantee that the voices generated are sincere.
In view of this, our project was aimed at developing TTS systems for creating natural voice outputs that take into consideration cultural differences and youthful appeal by building upon well-established neural network models in speech synthesis and by incorporating new techniques.
The results of our work are promising: we also managed to create a TTS model for accurately reading Urdu text, which was also marked to have perfect native-speaker-like pronunciation. These are the practical implications of our research across education, digital accessibility and media, possibly shifting popular culture. What we are trying to achieve is more friendly and natural biometrics for speech interfacing for Urdu users.

Keywords
Text-to-Speech
Urdu Language
Speech Synthesis
Computational Linguistics
Language Modeling
Natural Language Processing
Speech Processing
IoT-Based Smart Irrigation System Using Hybrid Ensemble Models for Water Usage Prediction
Using Convolutional Neural Networks for Enhanced Pneumonia Detection via Chest X-Rays