Abstract

Modern Text-to-Speech (TTS) systems demand massive training corpora, creating a severe bottleneck for low-resource languages like Polish. Furthermore, relying on commercial cloud-based TTS platforms introduces critical risks to data privacy and latency. This paper presents an effective cross-lingual fine-tuning strategy that adapts a pre-trained, English-centric non-autoregressive architecture (F5-TTS) for high-fidelity, locally deployable Polish speech synthesis. To overcome the open-source data deficit, we utilize a hybrid dataset combining crowdsourced amateur recordings with a proprietary professional studio corpus, which notably includes the deep and distinct voice of the prominent Polish actor Piotr Fronczewski. Objective metrics (UTMOS, WER) and subjective MUSHRA evaluations demonstrate that our fine-tuned model significantly outperforms the current local state-of-the-art baseline (XTTS-v2), achieving highly natural, artifact-free zero-shot voice cloning. Coupled with a lightweight on-premise GUI, this work provides a robust solution for Polish speech synthesis, serving as a blueprint for adapting a foundational TTS architecture to other underrepresented languages.

Recommended Citation

Szołkowski, B., Paszko, B., Podoba, Ł. & Szklanny, K.(2026). Adapting a Multilingual Zero-Shot Text-to-Speech Model for High-Fidelity Synthesis in Polish Despite Limited Training Data. In M. Valenta, B. Mannová, R. Pergl, A. Przybylek, M. Lang, H. Linger, C. Schneider, N. Iivari, & E. Insfran (Eds.), Making ISD Sustainable: Reloaded with AI and Automation (ISD2026 Proceedings). Prague, Czech Republic: Czech Technical University in Prague. ISBN: 978-80-01-07585-2. https://doi.org/10.62036/ISD.2026.39

Paper Type

Short Paper

DOI

10.62036/ISD.2026.39

Share

COinS
 

Adapting a Multilingual Zero-Shot Text-to-Speech Model for High-Fidelity Synthesis in Polish Despite Limited Training Data

Modern Text-to-Speech (TTS) systems demand massive training corpora, creating a severe bottleneck for low-resource languages like Polish. Furthermore, relying on commercial cloud-based TTS platforms introduces critical risks to data privacy and latency. This paper presents an effective cross-lingual fine-tuning strategy that adapts a pre-trained, English-centric non-autoregressive architecture (F5-TTS) for high-fidelity, locally deployable Polish speech synthesis. To overcome the open-source data deficit, we utilize a hybrid dataset combining crowdsourced amateur recordings with a proprietary professional studio corpus, which notably includes the deep and distinct voice of the prominent Polish actor Piotr Fronczewski. Objective metrics (UTMOS, WER) and subjective MUSHRA evaluations demonstrate that our fine-tuned model significantly outperforms the current local state-of-the-art baseline (XTTS-v2), achieving highly natural, artifact-free zero-shot voice cloning. Coupled with a lightweight on-premise GUI, this work provides a robust solution for Polish speech synthesis, serving as a blueprint for adapting a foundational TTS architecture to other underrepresented languages.