Abstract

The growing availability of data distributed across multiple independent sources creates new challenges for classification tasks. Such datasets often differ not only in their representations but also in class proportions, which complicates the construction of a consistent global model. This work addresses these issues by extending the authors’ previously proposed distributed classification framework that combines conflict analysis, coalition formation, and rule induction. The novelty is achieved by introducing a class balancing stage performed locally on each dataset prior to the learning process. Six data-level balancing techniques are evaluated. Decision rules are generated using four rough set-based algorithms, and final decisions are determined by three strategies. Experiments are conducted on two datasets: Car Evaluation and Balance Scale. The proposed approach outperforms a baseline for Car Evaluation, while for Balance Scale its effectiveness depends on data fragmentation. Additionally, coalition analysis indicates that higher data distribution leads to larger coalition sizes.

Recommended Citation

Kusztal, K. & Przybyła-Kasperek, M.(2026). Impact of Class Balancing on Coalition-Based Classification in Distributed Data Settings. In M. Valenta, B. Mannová, R. Pergl, A. Przybylek, M. Lang, H. Linger, C. Schneider, N. Iivari, & E. Insfran (Eds.), Making ISD Sustainable: Reloaded with AI and Automation (ISD2026 Proceedings). Prague, Czech Republic: Czech Technical University in Prague. ISBN: 978-80-01-07585-2. https://doi.org/10.62036/ISD.2026.71

Paper Type

Short Paper

DOI

10.62036/ISD.2026.71

Share

COinS
 

Impact of Class Balancing on Coalition-Based Classification in Distributed Data Settings

The growing availability of data distributed across multiple independent sources creates new challenges for classification tasks. Such datasets often differ not only in their representations but also in class proportions, which complicates the construction of a consistent global model. This work addresses these issues by extending the authors’ previously proposed distributed classification framework that combines conflict analysis, coalition formation, and rule induction. The novelty is achieved by introducing a class balancing stage performed locally on each dataset prior to the learning process. Six data-level balancing techniques are evaluated. Decision rules are generated using four rough set-based algorithms, and final decisions are determined by three strategies. Experiments are conducted on two datasets: Car Evaluation and Balance Scale. The proposed approach outperforms a baseline for Car Evaluation, while for Balance Scale its effectiveness depends on data fragmentation. Additionally, coalition analysis indicates that higher data distribution leads to larger coalition sizes.