Paper Type
Complete
Paper Number
PACIS2026-1210
Description
Duplicate and near-duplicate submissions (e.g., ideas) pose a persistent challenge in crowdsourcing and idea management platforms (IMPs), inflating evaluation effort, fragmenting attention, and obscuring novel contributions. This study investigates how large language model (LLM)-based semantic similarity detection can support idea screening in organizational innovation settings. Following a design science research approach, we developed and evaluated an operational platform artifact that integrates embedding-based cosine similarity and LLM-assisted classification into a human-in-the-loop workflow. Using two real-world datasets from a higher education institution – idea submissions (i.e., short text data) and project proposals (i.e., long text data) – we conducted pairwise comparisons to assess detection accuracy and threshold calibration. The results show that similarity detection is substantially more reliable for longer and structured texts, while near-duplicates remain difficult to identify and require human oversight. This study contributes empirical evidence on LLM-based idea screening in IMPs and offers design guidance for responsible deployment.
Recommended Citation
Simic, Dejan and Leible, Stephan, "Examining Idea Screening Through Large Language Model-Based Similarity Detection in Idea Management Platforms" (2026). PACIS 2026 Proceedings. 3.
https://aisel.aisnet.org/pacis2026/ai_ml/ai_ml/3
Examining Idea Screening Through Large Language Model-Based Similarity Detection in Idea Management Platforms
Duplicate and near-duplicate submissions (e.g., ideas) pose a persistent challenge in crowdsourcing and idea management platforms (IMPs), inflating evaluation effort, fragmenting attention, and obscuring novel contributions. This study investigates how large language model (LLM)-based semantic similarity detection can support idea screening in organizational innovation settings. Following a design science research approach, we developed and evaluated an operational platform artifact that integrates embedding-based cosine similarity and LLM-assisted classification into a human-in-the-loop workflow. Using two real-world datasets from a higher education institution – idea submissions (i.e., short text data) and project proposals (i.e., long text data) – we conducted pairwise comparisons to assess detection accuracy and threshold calibration. The results show that similarity detection is substantially more reliable for longer and structured texts, while near-duplicates remain difficult to identify and require human oversight. This study contributes empirical evidence on LLM-based idea screening in IMPs and offers design guidance for responsible deployment.
Comments
01-AIML