Abstract

Through the infrastructure of our partner Next Mobile P.S.A. around 150-170 million unfiltered connection attempts are transferred every month. The data of 3.3 million phone numbers from August 2024 was divided into 21 classes. Then, the research used several machine learning and artificial intelligence methods, e.g. k-means, LSTM, SVM, decision tree (Fine Tree), optimized decision tree, Gaussian Naive Bayes, KNN, and Random Forest. The best approach was to use in the first step unsupervised methods, then compare the numbers with the online opinion even if it is limited to around 200 thousand numbers, and use supervised methods at the end with classes as cluster numbers. We have received the accuracy around 99.4% with true positive and true negative rates 99% as well for the chosen group numbers. Our final solution, which is a blend of machine learning and statistical analysis, is the final tool to spot suspicious numbers.

Recommended Citation

Hłobaż, A. & Milczarski, P.(2026). Vishing Detection using Statistical Analysis and Machine Learning Methods. In M. Valenta, B. Mannová, R. Pergl, A. Przybylek, M. Lang, H. Linger, C. Schneider, N. Iivari, & E. Insfran (Eds.), Making ISD Sustainable: Reloaded with AI and Automation (ISD2026 Proceedings). Prague, Czech Republic: Czech Technical University in Prague. ISBN: 978-80-01-07585-2. https://doi.org/10.62036/ISD.2026.207

Paper Type

Poster

DOI

10.62036/ISD.2026.207

Share

COinS
 

Vishing Detection using Statistical Analysis and Machine Learning Methods

Through the infrastructure of our partner Next Mobile P.S.A. around 150-170 million unfiltered connection attempts are transferred every month. The data of 3.3 million phone numbers from August 2024 was divided into 21 classes. Then, the research used several machine learning and artificial intelligence methods, e.g. k-means, LSTM, SVM, decision tree (Fine Tree), optimized decision tree, Gaussian Naive Bayes, KNN, and Random Forest. The best approach was to use in the first step unsupervised methods, then compare the numbers with the online opinion even if it is limited to around 200 thousand numbers, and use supervised methods at the end with classes as cluster numbers. We have received the accuracy around 99.4% with true positive and true negative rates 99% as well for the chosen group numbers. Our final solution, which is a blend of machine learning and statistical analysis, is the final tool to spot suspicious numbers.