Abstract
Data quality remains a critical challenge in data-centric information systems, particularly in the presence of high-dimensional data, class imbalance, and limited labeled data. Undetected outliers may propagate through analytical pipelines and reduce the reliability of downstream decision-support services. This paper proposes a hybrid outlier detection framework that combines heterogeneous unsupervised detectors through decision-level majority voting to generate pseudo-labels, followed by supervised refinement in a reduced-dimensional feature space. The framework investigates how detector complementarity affects pseudo-label reliability and downstream model performance. Experiments on ADBench datasets show that detector diversity and partial error decorrelation improve detection robustness. The proposed approach outperforms standalone detectors and identifies conditions under which supervised refinement improves or degrades detection quality. By framing outlier detection as a modular validation layer, the framework supports the development of more reliable and deployable data-centric information systems.
Paper Type
Short Paper
DOI
10.62036/ISD.2026.70
Hybrid Outlier Detection via Majority-Vote Pseudo-Labeling
Data quality remains a critical challenge in data-centric information systems, particularly in the presence of high-dimensional data, class imbalance, and limited labeled data. Undetected outliers may propagate through analytical pipelines and reduce the reliability of downstream decision-support services. This paper proposes a hybrid outlier detection framework that combines heterogeneous unsupervised detectors through decision-level majority voting to generate pseudo-labels, followed by supervised refinement in a reduced-dimensional feature space. The framework investigates how detector complementarity affects pseudo-label reliability and downstream model performance. Experiments on ADBench datasets show that detector diversity and partial error decorrelation improve detection robustness. The proposed approach outperforms standalone detectors and identifies conditions under which supervised refinement improves or degrades detection quality. By framing outlier detection as a modular validation layer, the framework supports the development of more reliable and deployable data-centric information systems.
Recommended Citation
Duraj, A. & Warchoł, F.(2026). Hybrid Outlier Detection via Majority-Vote Pseudo-Labeling. In M. Valenta, B. Mannová, R. Pergl, A. Przybylek, M. Lang, H. Linger, C. Schneider, N. Iivari, & E. Insfran (Eds.), Making ISD Sustainable: Reloaded with AI and Automation (ISD2026 Proceedings). Prague, Czech Republic: Czech Technical University in Prague. ISBN: 978-80-01-07585-2. https://doi.org/10.62036/ISD.2026.70