Abstract
Smishing--phishing conducted via SMS--continues to spread globally. Most detection systems are built using only English data, which limits their use for other languages. We propose a multilingual smishing detection framework based on self-training with XLM-RoBERTa. Starting with English-labeled data, we apply zero-shot inference to Bengali and Swahili messages, extract high-confidence predictions, and incorporate them as pseudo-labeled examples for fine-tuning. Our approach improves recall and F1 scores in both target languages without requiring parallel corpora or costly annotations. We further analyze the impact of language-specific and balanced pseudo-label augmentation through ablation studies. The results show that the combination of samples from multiple languages leads to better generalization and reliability. This work highlights an effective low-resource strategy for building multilingual smishing detectors, allowing a broader deployment across linguistically diverse user populations.
Paper Type
Short Paper
DOI
10.62036/ISD.2026.202
Self-Training Approach for Smishing Detection in Multilingual and Low-Resource Settings
Smishing--phishing conducted via SMS--continues to spread globally. Most detection systems are built using only English data, which limits their use for other languages. We propose a multilingual smishing detection framework based on self-training with XLM-RoBERTa. Starting with English-labeled data, we apply zero-shot inference to Bengali and Swahili messages, extract high-confidence predictions, and incorporate them as pseudo-labeled examples for fine-tuning. Our approach improves recall and F1 scores in both target languages without requiring parallel corpora or costly annotations. We further analyze the impact of language-specific and balanced pseudo-label augmentation through ablation studies. The results show that the combination of samples from multiple languages leads to better generalization and reliability. This work highlights an effective low-resource strategy for building multilingual smishing detectors, allowing a broader deployment across linguistically diverse user populations.
Recommended Citation
Krawczyk, N., Probierz, B., Olszak, C., Żurada, J., Hatami, Z. & Kozak, J.(2026). Self-Training Approach for Smishing Detection in Multilingual and Low-Resource Settings. In M. Valenta, B. Mannová, R. Pergl, A. Przybylek, M. Lang, H. Linger, C. Schneider, N. Iivari, & E. Insfran (Eds.), Making ISD Sustainable: Reloaded with AI and Automation (ISD2026 Proceedings). Prague, Czech Republic: Czech Technical University in Prague. ISBN: 978-80-01-07585-2. https://doi.org/10.62036/ISD.2026.202