Abstract
Large language models (LLMs) are typically deployed as standalone systems, although many professional tasks require structured collaboration among specialists. This paper presents a proof-of-concept design and preliminary evaluation of a multi-agent LLM framework for multidisciplinary case deliberation, using a simulated medical board as a high-stakes test scenario. A Python-based orchestrator manages specialist roles and agent communication, aggregating their findings into a structured clinical assessment. We apply the framework to four literature-based clinical cases and compare its outputs against reference clinical pathways using embedding-based semantic similarity across six categories: confirmed diagnoses, suspected diagnoses, treatment plans, risks, next steps, and notes. We further compare the multi-agent setup with a single-model baseline. The preliminary results indicate that the multi-agent approach achieves stronger alignment in several clinically relevant categories. Especially when multi-perspective reasoning and risk synthesis is required. Although the advantage is not uniform and the study is subject to important limitations: a very small case set, absence of clinician evaluation, potential training-data contamination from published sources, and the use of an auxiliary LLM for reference extraction. Limited generalization requires further expert-led validation, but the study proves that orchestrating specialized LLMs for expert-style deliberation is feasible.
Paper Type
Short Paper
DOI
10.62036/ISD.2026.66
From Independent LLMs to Collaborative Networks: Implementation and Evaluation of a Multi-Agent Medical Board
Large language models (LLMs) are typically deployed as standalone systems, although many professional tasks require structured collaboration among specialists. This paper presents a proof-of-concept design and preliminary evaluation of a multi-agent LLM framework for multidisciplinary case deliberation, using a simulated medical board as a high-stakes test scenario. A Python-based orchestrator manages specialist roles and agent communication, aggregating their findings into a structured clinical assessment. We apply the framework to four literature-based clinical cases and compare its outputs against reference clinical pathways using embedding-based semantic similarity across six categories: confirmed diagnoses, suspected diagnoses, treatment plans, risks, next steps, and notes. We further compare the multi-agent setup with a single-model baseline. The preliminary results indicate that the multi-agent approach achieves stronger alignment in several clinically relevant categories. Especially when multi-perspective reasoning and risk synthesis is required. Although the advantage is not uniform and the study is subject to important limitations: a very small case set, absence of clinician evaluation, potential training-data contamination from published sources, and the use of an auxiliary LLM for reference extraction. Limited generalization requires further expert-led validation, but the study proves that orchestrating specialized LLMs for expert-style deliberation is feasible.
Recommended Citation
Czyzewski, A., Szczodrak, M. & Zielonka, M.(2026). From Independent LLMs to Collaborative Networks: Implementation and Evaluation of a Multi-Agent Medical Board. In M. Valenta, B. Mannová, R. Pergl, A. Przybylek, M. Lang, H. Linger, C. Schneider, N. Iivari, & E. Insfran (Eds.), Making ISD Sustainable: Reloaded with AI and Automation (ISD2026 Proceedings). Prague, Czech Republic: Czech Technical University in Prague. ISBN: 978-80-01-07585-2. https://doi.org/10.62036/ISD.2026.66