IRAIS 2026 Proceedings

Abstract

Internet of Things (IoT) networks expose organizations to attacks that were absent from a detector's training data. A closed-set classifier must still assign such traffic to a known class, even when the label is not meaningful (Geng et al., 2021). Open-set recognition addresses this problem by allowing rejection as unknown (Scheirer et al., 2013). The Extreme Value Machine (EVM) models the smallest half-distances from each positive reference sample to nearby negative-class samples with Weibull distributions and converts those margins into class-inclusion scores (Rudd et al., 2018). A sample is accepted as known only when its largest inclusion score reaches a chosen threshold. Raising that threshold flags more traffic as unknown, which may catch more novel attacks but also creates more alerts. The threshold is therefore an operating decision, not simply a value to leave at a software default.

We ask how an organization should govern this operating point when missed attacks and false alarms have unequal consequences and analyst attention is limited. Signal detection theory (SDT) separates sensitivity, the detector's ability to distinguish signal from noise, from criterion, the cutoff used to act on that evidence (Green & Swets, 1966). The EVM representation and fitted margins determine sensitivity; the inclusion threshold sets the criterion. Changing the threshold can move unknown recall and false alarms without changing the detector's underlying sensitivity. We define threshold governance as the process through which security managers authorize, document, monitor, and revise that criterion using two inputs outside the model: the relative weight placed on a missed novel attack and the number of alerts the security operations center (SOC) can review. Cyber-loss evidence shows that a small set of extreme events can be very costly, but loss severity also varies by event and organization (Eling & Wirfs, 2019). High false-alarm rates can also reduce analyst performance (Layman & Roden, 2023). These facts make loss weight and review capacity local governance inputs, not universal constants.

The argument yields three propositions. P1 states that, at a matched known-traffic false-positive rate, a correctly implemented EVM will produce higher unknown-attack recall than global-score baselines across held-out attack families. This expectation can fail if per-reference tail calibration offers no advantage in the observed feature space. P2 states that increasing the effective weight on missed attacks will move the selected criterion toward more liberal unknown flagging, although the size of the move depends on the empirical operating curve. It can fail when the same operating point remains optimal across the tested weights. P3 states that a review budget will produce a less liberal, capacity-feasible criterion when the loss-minimizing point creates more alerts than the SOC can examine. It can fail when the budget never binds. P1 is a technical comparison; P2 and P3 concern the governed choice among technically available operating points.

Prior work has applied EVM to IoT intrusion detection (Safari & Kim, 2025). We propose a computational policy simulation using Kitsune (Mirsky et al., 2018) and IoT-23 (García et al., 2020). Each test will hold out an entire attack family as unknown. Malicious flows from that family will be the unknown target, while benign flows will remain known traffic; treating all flows from a capture as unknown would inflate recall and alert counts. Feature selection, scaling, model fitting, and criterion calibration will use only known-class training and validation data. The untouched test partition will compare a reference implementation of the EVM with Isolation Forest and One-Class SVM at the same validation-set false-positive targets. We will report unknown recall, unknown precision, known false-positive rate, alert count, and variation across held-out families and repeated splits, because one split cannot establish stable predictive performance (Shmueli & Koppius, 2011). For P2, miss weights from 1 to 1,000 are sensitivity-analysis assumptions that combine attack frequency and relative loss; they are not estimated dollar costs. For P3, review budgets of 50, 500, and 5,000 alerts per evaluation window are explicit simulation parameters, not measured daily SOC capacity.

This design separates what the benchmark can establish from what still requires organizational evidence. The simulation can identify a detection-and-alert frontier, show whether a loss-weighted operating point changes, and show when an assumed review budget binds. It cannot show that SOC managers will choose that point, that the assumed weights equal realized cyber losses, or that an alert budget matches analyst capacity in practice. Both datasets are controlled IoT benchmarks, so they also cannot establish effects on analyst behavior or organizational effectiveness. The contribution is a bounded threshold governance logic: estimate the local operating curve, state the loss assumptions, select a criterion within an explicit review budget, and require reauthorization when traffic, threats, or staffing change. Field studies with SOC managers and operational logs are needed to test the organizational propositions and replace simulated budgets with observed capacity.

References

Eling, M., & Wirfs, J. (2019). What are the actual costs of cyber risk events? European Journal of Operational Research, 272(3), 1109–1119.

García, S., Parmisano, A., & Erquiaga, M. J. (2020). IoT-23: A labeled dataset with malicious and benign IoT network traffic (Version 1.0.0) [Data set]. Zenodo.

Geng, C., Huang, S.-J., & Chen, S. (2021). Recent advances in open set recognition: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(10), 3614–3631.

Green, D. M., & Swets, J. A. (1966). Signal detection theory and psychophysics. Wiley.

Layman, L., & Roden, W. (2023). A controlled experiment on the impact of intrusion detection false alarm rate on analyst performance. Proceedings of the Human Factors and Ergonomics Society Annual Meeting, 67(1), 220–225.

Mirsky, Y., Doitshman, T., Elovici, Y., & Shabtai, A. (2018). Kitsune: An ensemble of autoencoders for online network intrusion detection. Proceedings of the 2018 Network and Distributed System Security Symposium (NDSS).

Rudd, E. M., Jain, L. P., Scheirer, W. J., & Boult, T. E. (2018). The extreme value machine. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(3), 762–768.

Safari, A., & Kim, D. J. (2025). Enhancing IoT security and information systems resilience: An extreme value machine approach. ICIS 2025 Proceedings, 22. https://aisel.aisnet.org/icis2025/cyb_security/cyb_security/22

Scheirer, W. J., de Rezende Rocha, A., Sapkota, A., & Boult, T. E. (2013). Toward open set recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(7), 1757–1772.

Shmueli, G., & Koppius, O. R. (2011). Predictive analytics in information systems research. MIS Quarterly, 35(3), 553–572.

Abstract Only

Share

COinS
 
 

To view the content in your browser, please download Adobe Reader or, alternately,
you may Download the file to your hard drive.

NOTE: The latest versions of Adobe Reader do not support viewing PDF files within Firefox on Mac OS and if you are using a modern (Intel) Mac, there is no official plugin for viewing PDF files within the browser window.