Paper Type
Short
Paper Number
PACIS2026-1978
Description
Large Language Models (LLMs) show promise for open-ended medical diagnosis, yet single-model predictions often suffer from instability, inconsistency, and limited robustness in complex clinical cases. To address this challenge, this study proposes Probabilistic Multi-LLM Fusion (PMLF), an unsupervised collective intelligence framework that aggregates diagnostic outputs from multiple LLMs at the probability-distribution level. The framework first standardizes model-generated diagnoses into a unified medical terminology space and calibrates confidence scores into comparable probability distributions. It then infers a latent consensus diagnostic distribution while jointly modeling model competence and case difficulty. Using 2,497 real-world clinical cases, experimental results show that PMLF consistently outperforms mean probability averaging, majority voting, and frequency-based aggregation across Top-K accuracy, mean reciprocal rank, and coverage. The findings demonstrate the potential of probabilistic collective intelligence to improve the reliability and robustness of LLM-assisted open-ended medical diagnosis.
Recommended Citation
Yan, Zhijun; ZHU, BO; Zhang, Zhu; Tang, Mingrui; and Wang, Tianmei, "LLM Collective Intelligence for Open-Ended Medical Diagnosis" (2026). PACIS 2026 Proceedings. 22.
https://aisel.aisnet.org/pacis2026/ishealthcare/ishealthcare/22
LLM Collective Intelligence for Open-Ended Medical Diagnosis
Large Language Models (LLMs) show promise for open-ended medical diagnosis, yet single-model predictions often suffer from instability, inconsistency, and limited robustness in complex clinical cases. To address this challenge, this study proposes Probabilistic Multi-LLM Fusion (PMLF), an unsupervised collective intelligence framework that aggregates diagnostic outputs from multiple LLMs at the probability-distribution level. The framework first standardizes model-generated diagnoses into a unified medical terminology space and calibrates confidence scores into comparable probability distributions. It then infers a latent consensus diagnostic distribution while jointly modeling model competence and case difficulty. Using 2,497 real-world clinical cases, experimental results show that PMLF consistently outperforms mean probability averaging, majority voting, and frequency-based aggregation across Top-K accuracy, mean reciprocal rank, and coverage. The findings demonstrate the potential of probabilistic collective intelligence to improve the reliability and robustness of LLM-assisted open-ended medical diagnosis.
Comments
14-Healthcare