Paper Type
Complete
Abstract
As information systems increasingly adapt to users’ domain knowledge, scalable methods are needed to infer knowledge from naturally occurring explanations. Yet the capacity of unsupervised short-text analytics to differentiate users by underlying knowledge remains underexamined. This study investigates whether brief, open-ended domain explanations can support unsupervised user segmentation and how text representation influences segmentation outcomes. Using cybersecurity as a case study, we analyze responses from 358 participants explaining HTTPS and authentication mechanisms. We apply K-means clustering to multiple text representation methods and evaluate resulting segments post hoc using objective and self-reported cybersecurity knowledge measures. Results provide evidence that unsupervised clustering can differentiate users from short-text explanations, with domain-adapted and lexical representations exhibiting the strongest alignment with independent knowledge measures, while general-purpose embeddings were less effective. The study contributes a representation-sensitive approach to scalable knowledge segmentation, suggesting that knowledge differentiation in unsupervised text analytics depends on representation design.
Paper Number
1276
Recommended Citation
Cotoranu, Andreea and Chen, Li-Chiou, "Segmenting Cybersecurity Knowledge from Short Explanations: An Unsupervised Analytics Approach" (2026). AMCIS 2026 Proceedings. 7.
https://aisel.aisnet.org/amcis2026/sig_dsa/sig_dsa/7
Segmenting Cybersecurity Knowledge from Short Explanations: An Unsupervised Analytics Approach
As information systems increasingly adapt to users’ domain knowledge, scalable methods are needed to infer knowledge from naturally occurring explanations. Yet the capacity of unsupervised short-text analytics to differentiate users by underlying knowledge remains underexamined. This study investigates whether brief, open-ended domain explanations can support unsupervised user segmentation and how text representation influences segmentation outcomes. Using cybersecurity as a case study, we analyze responses from 358 participants explaining HTTPS and authentication mechanisms. We apply K-means clustering to multiple text representation methods and evaluate resulting segments post hoc using objective and self-reported cybersecurity knowledge measures. Results provide evidence that unsupervised clustering can differentiate users from short-text explanations, with domain-adapted and lexical representations exhibiting the strongest alignment with independent knowledge measures, while general-purpose embeddings were less effective. The study contributes a representation-sensitive approach to scalable knowledge segmentation, suggesting that knowledge differentiation in unsupervised text analytics depends on representation design.
When commenting on articles, please be friendly, welcoming, respectful and abide by the AIS eLibrary Discussion Thread Code of Conduct posted here.
Comments
SIG DSA