Paper Type

Short

Paper Number

PACIS2026-1978

Description

Large Language Models (LLMs) show promise for open-ended medical diagnosis, yet single-model predictions often suffer from instability, inconsistency, and limited robustness in complex clinical cases. To address this challenge, this study proposes Probabilistic Multi-LLM Fusion (PMLF), an unsupervised collective intelligence framework that aggregates diagnostic outputs from multiple LLMs at the probability-distribution level. The framework first standardizes model-generated diagnoses into a unified medical terminology space and calibrates confidence scores into comparable probability distributions. It then infers a latent consensus diagnostic distribution while jointly modeling model competence and case difficulty. Using 2,497 real-world clinical cases, experimental results show that PMLF consistently outperforms mean probability averaging, majority voting, and frequency-based aggregation across Top-K accuracy, mean reciprocal rank, and coverage. The findings demonstrate the potential of probabilistic collective intelligence to improve the reliability and robustness of LLM-assisted open-ended medical diagnosis.

Comments

14-Healthcare

Share

COinS
 
Jul 5th, 12:00 AM

LLM Collective Intelligence for Open-Ended Medical Diagnosis

Large Language Models (LLMs) show promise for open-ended medical diagnosis, yet single-model predictions often suffer from instability, inconsistency, and limited robustness in complex clinical cases. To address this challenge, this study proposes Probabilistic Multi-LLM Fusion (PMLF), an unsupervised collective intelligence framework that aggregates diagnostic outputs from multiple LLMs at the probability-distribution level. The framework first standardizes model-generated diagnoses into a unified medical terminology space and calibrates confidence scores into comparable probability distributions. It then infers a latent consensus diagnostic distribution while jointly modeling model competence and case difficulty. Using 2,497 real-world clinical cases, experimental results show that PMLF consistently outperforms mean probability averaging, majority voting, and frequency-based aggregation across Top-K accuracy, mean reciprocal rank, and coverage. The findings demonstrate the potential of probabilistic collective intelligence to improve the reliability and robustness of LLM-assisted open-ended medical diagnosis.