Paper Type

Complete

Paper Number

PACIS2026-1863

Description

Teamwork is crucial for systematic literature reviews (SLRs). This paper aims to develop a method that marries bibliometry with embedding models to simulate the cooperation of complementary SLR team members to improve and automate screening. The study analyzed all 1,023 subsets of 10 embedding models from five providers across 597 Scopus records, using 10 bibliometric indicators. Analysis of models' consensus, union, and average results shows that the top five-model union combination achieves a normalized score of 2.232 with 95 unique articles - covering 137.5% more ground compared to the best single model. full consensus drops from 40 articles (k=1) to 3 (k=10), but consensus articles demonstrate stronger bibliometric coherence (keyword coherence B=1.333 at k=10, roughly three times the single model's average). Pairwise Jaccard analysis reveals high redundancy among providers (OpenAI pair J=0.569) and high diversity across providers (minimum J=0.176), indicating that using 2–5 diverse models yields the best results.

Comments

15-Method

Share

COinS
 
Jul 5th, 12:00 AM

Optimal Multi-Model Embedding Combinations for Bibliometric Screening in Systematic Literature Reviews

Teamwork is crucial for systematic literature reviews (SLRs). This paper aims to develop a method that marries bibliometry with embedding models to simulate the cooperation of complementary SLR team members to improve and automate screening. The study analyzed all 1,023 subsets of 10 embedding models from five providers across 597 Scopus records, using 10 bibliometric indicators. Analysis of models' consensus, union, and average results shows that the top five-model union combination achieves a normalized score of 2.232 with 95 unique articles - covering 137.5% more ground compared to the best single model. full consensus drops from 40 articles (k=1) to 3 (k=10), but consensus articles demonstrate stronger bibliometric coherence (keyword coherence B=1.333 at k=10, roughly three times the single model's average). Pairwise Jaccard analysis reveals high redundancy among providers (OpenAI pair J=0.569) and high diversity across providers (minimum J=0.176), indicating that using 2–5 diverse models yields the best results.