Abstract

Healthcare organisations face a critical gap in AI evaluation: accuracy benchmarks do not indicate whether systems will function reliably within the constraints of real clinical environments. This study addresses that gap with two contributions. First, it introduces the enrollability framework, four Information Systems-level criteria, grounded in Actor-Network Theory, that define when a Natural Language Processing (NLP) method can be stabilised as a clinical infrastructure. Second, it examines how Large Language Models (LLMs) align with these requirements in a non-English clinical NLP setting. While rule-based systems encode stable inscriptions compatible with network expectations, LLMs rely on next-token prediction that mimics but does not ensure such stability. We empirically compare rule-based and LLM-based approaches (Llama-3-8B-it, Gemma-7b-it) on 1,679 Polish paediatric epicrises. Findings reveal that accuracy and enrollability yield divergent evaluations, and suggest that common LLM failures (e.g., hallucinations, inter-physician variability, and notation mismatch) may be associated with properties of their generative mechanism, beyond capability limits.

Recommended Citation

Tworek, P., Khan, Y., Bargieł, M., Pełech-Pilichowski, T., Mikołajczyk, M., Lewandowski, R. & Sousa, J.(2026). Actor-Network Theory Analysis of When LLMs Can and Cannot Be Enrolled. In M. Valenta, B. Mannová, R. Pergl, A. Przybylek, M. Lang, H. Linger, C. Schneider, N. Iivari, & E. Insfran (Eds.), Making ISD Sustainable: Reloaded with AI and Automation (ISD2026 Proceedings). Prague, Czech Republic: Czech Technical University in Prague. ISBN: 978-80-01-07585-2. https://doi.org/10.62036/ISD.2026.110

Paper Type

Short Paper

DOI

10.62036/ISD.2026.110

Share

COinS
 

Actor-Network Theory Analysis of When LLMs Can and Cannot Be Enrolled

Healthcare organisations face a critical gap in AI evaluation: accuracy benchmarks do not indicate whether systems will function reliably within the constraints of real clinical environments. This study addresses that gap with two contributions. First, it introduces the enrollability framework, four Information Systems-level criteria, grounded in Actor-Network Theory, that define when a Natural Language Processing (NLP) method can be stabilised as a clinical infrastructure. Second, it examines how Large Language Models (LLMs) align with these requirements in a non-English clinical NLP setting. While rule-based systems encode stable inscriptions compatible with network expectations, LLMs rely on next-token prediction that mimics but does not ensure such stability. We empirically compare rule-based and LLM-based approaches (Llama-3-8B-it, Gemma-7b-it) on 1,679 Polish paediatric epicrises. Findings reveal that accuracy and enrollability yield divergent evaluations, and suggest that common LLM failures (e.g., hallucinations, inter-physician variability, and notation mismatch) may be associated with properties of their generative mechanism, beyond capability limits.