Abstract
This project investigates the use of self-hosted Large Language Models (LLMs) for processing Polish medical text within a GDPR-compliant, on-prem environment. The goal is to assess whether relatively small, locally deployed LLMs can be used in tasks such as data extraction and translation of Polish medical records. We tested our solution on anonymized pediatric data set of 10 years of activity in Tuchola Hospital (Szpital Tucholski). We evaluated the performance of small and medium quantized models on translation and extraction tasks, measured the energy efficiency, and validated the quality of model generated output against ground truth. Experiments revealed that translation and extraction quality depend on the quality of the source text, and using the biggest model possible is not cost-optimal in terms of used energy vs. end accuracy, with smaller models with extra context performing up to 7pp. better than bigger models for translation and 25pp. for extraction.
Paper Type
Short Paper
DOI
10.62036/ISD.2026.201
Self-hosted Large Language Models for Medical Text Data Mining
This project investigates the use of self-hosted Large Language Models (LLMs) for processing Polish medical text within a GDPR-compliant, on-prem environment. The goal is to assess whether relatively small, locally deployed LLMs can be used in tasks such as data extraction and translation of Polish medical records. We tested our solution on anonymized pediatric data set of 10 years of activity in Tuchola Hospital (Szpital Tucholski). We evaluated the performance of small and medium quantized models on translation and extraction tasks, measured the energy efficiency, and validated the quality of model generated output against ground truth. Experiments revealed that translation and extraction quality depend on the quality of the source text, and using the biggest model possible is not cost-optimal in terms of used energy vs. end accuracy, with smaller models with extra context performing up to 7pp. better than bigger models for translation and 25pp. for extraction.
Recommended Citation
Katulski, F., Tworek, P., Chrobot, A., Baliś, B. & Sousa, J.(2026). Self-hosted Large Language Models for Medical Text Data Mining. In M. Valenta, B. Mannová, R. Pergl, A. Przybylek, M. Lang, H. Linger, C. Schneider, N. Iivari, & E. Insfran (Eds.), Making ISD Sustainable: Reloaded with AI and Automation (ISD2026 Proceedings). Prague, Czech Republic: Czech Technical University in Prague. ISBN: 978-80-01-07585-2. https://doi.org/10.62036/ISD.2026.201