Abstract

Efficient fine-tuning of an automatic speech recognition system is a key challenge in machine learning research. A major obstacle is accurately recognizing jargon-heavy, domain-specific speech in expert discussions, such as those among lawyers, doctors, or engineers. We propose an automated machine-learning approach that fine-tunes the Whisper small model to enhance medical speech recognition. This fine-tuning uses three selected PEFT methods (LoRA, DoRA, and MoRA) and employs a multi-objective optimization framework, enabling a balance between the accuracy of medical speech recognition (evaluated on the ADMEDVOICE test subset) and general Polish speech recognition (evaluated on the Common Voice test subset). The set of Pareto-optimal solutions was found to be statistically significantly better than those obtained from randomly chosen configurations and to provide better general language performance than most solutions generated by the single-objective optimization process. Solutions on this front are competitive with state-of-the-art systems for speech recognition in the medical domain.

Recommended Citation

Kurowski, A., Zaporowski, S., Kostek, B. & Czyżewski, A.(2026). A multi-objective parameter-efficient fine-tuning method for medical Polish language automatic speech transcription systems. In M. Valenta, B. Mannová, R. Pergl, A. Przybylek, M. Lang, H. Linger, C. Schneider, N. Iivari, & E. Insfran (Eds.), Making ISD Sustainable: Reloaded with AI and Automation (ISD2026 Proceedings). Prague, Czech Republic: Czech Technical University in Prague. ISBN: 978-80-01-07585-2. https://doi.org/10.62036/ISD.2026.36

Paper Type

Full Paper

DOI

10.62036/ISD.2026.36

Share

COinS
 

A multi-objective parameter-efficient fine-tuning method for medical Polish language automatic speech transcription systems

Efficient fine-tuning of an automatic speech recognition system is a key challenge in machine learning research. A major obstacle is accurately recognizing jargon-heavy, domain-specific speech in expert discussions, such as those among lawyers, doctors, or engineers. We propose an automated machine-learning approach that fine-tunes the Whisper small model to enhance medical speech recognition. This fine-tuning uses three selected PEFT methods (LoRA, DoRA, and MoRA) and employs a multi-objective optimization framework, enabling a balance between the accuracy of medical speech recognition (evaluated on the ADMEDVOICE test subset) and general Polish speech recognition (evaluated on the Common Voice test subset). The set of Pareto-optimal solutions was found to be statistically significantly better than those obtained from randomly chosen configurations and to provide better general language performance than most solutions generated by the single-objective optimization process. Solutions on this front are competitive with state-of-the-art systems for speech recognition in the medical domain.