Abstract

Classifying natural language requirements (NLRs) is a challenging task. Recent studies show that Large Language Models (LLMs) can support automated categorisation of requirements; however, limited research has specifically addressed the distinction between Security Requirements (SRs) and Non-Security Requirements (NSRs). In this work, we investigate the reliability of Generative Pretrained Transformer (GPT)-like models, such as GPT-5.1, when classifying NLRs into SRs and NSRs by using a Zero-Shot Learning (ZSL) approach. Moreover, we explore how different industrial domains can affect the reliability of classification results. We evaluate the results by using standard machine learning evaluation metrics: F1-score (F1), Precision (P), and Recall (R). The model’s reliability varies considerably across application domains, achieving near-human performance in financial systems, moderate performance in security-standard domains, and markedly poorer performance in network-infrastructure domains. The findings highlight the need for caution in practice and the importance of developing more robust, domain-aware approaches.

Recommended Citation

Chatzipetrou, P., Gao, S. & Karlsson, F.(2026). Using Large Language Models to Classify Security Requirements: An Empirical Study. In M. Valenta, B. Mannová, R. Pergl, A. Przybylek, M. Lang, H. Linger, C. Schneider, N. Iivari, & E. Insfran (Eds.), Making ISD Sustainable: Reloaded with AI and Automation (ISD2026 Proceedings). Prague, Czech Republic: Czech Technical University in Prague. ISBN: 978-80-01-07585-2. https://doi.org/10.62036/ISD.2026.11

Paper Type

Short Paper

DOI

10.62036/ISD.2026.11

Share

COinS
 

Using Large Language Models to Classify Security Requirements: An Empirical Study

Classifying natural language requirements (NLRs) is a challenging task. Recent studies show that Large Language Models (LLMs) can support automated categorisation of requirements; however, limited research has specifically addressed the distinction between Security Requirements (SRs) and Non-Security Requirements (NSRs). In this work, we investigate the reliability of Generative Pretrained Transformer (GPT)-like models, such as GPT-5.1, when classifying NLRs into SRs and NSRs by using a Zero-Shot Learning (ZSL) approach. Moreover, we explore how different industrial domains can affect the reliability of classification results. We evaluate the results by using standard machine learning evaluation metrics: F1-score (F1), Precision (P), and Recall (R). The model’s reliability varies considerably across application domains, achieving near-human performance in financial systems, moderate performance in security-standard domains, and markedly poorer performance in network-infrastructure domains. The findings highlight the need for caution in practice and the importance of developing more robust, domain-aware approaches.