Bol (بول) - Urdu Text-to-Speech with Voice Cloning
A production Urdu TTS platform powered by a fine-tuned state-of-the-art speech model trained on 18,000+ native audio samples - delivering human-quality Urdu synthesis and zero-shot voice cloning through a polished web application. Rated 10/10 by native Urdu evaluators.
Key Outcome
10/10 naturalness rating
Rated indistinguishable from a native human voice by independent Urdu evaluators.
Over 170 million Urdu speakers worldwide lack access to natural-sounding text-to-speech. Existing TTS solutions produce robotic, unnatural Urdu - unusable for accessibility tools, content creation, or voice interfaces. No public model had been fine-tuned specifically on diverse, high-quality native Urdu speech data.
Fine-tuned a state-of-the-art speech model on a curated dataset of 18,000+ Urdu audio samples from native speakers. Implemented zero-shot voice cloning from a 10–30 second audio sample. Deployed as a full-stack web application with real-time audio visualisation, secure authentication, and generation history - achieving 10/10 naturalness ratings from native evaluators.
Key Features
Fine-tuned speech model on 18,000+ native Urdu audio samples
Zero-shot voice cloning from a 10–30 second WAV/MP3 sample
Real-time audio visualisation with animated 3D sphere player
Similarity control slider (stable → creative) for voice cloning
Generation history with playback and download per session
Secure user authentication and personal voice library
Dark mode support and clean bilingual (Urdu RTL / English) UI
10/10 naturalness rating from independent native Urdu evaluators
Technology Stack
Project Visuals




Interested in a similar project?
Tell us your requirements and we will respond within one business day.
