Location

Hilton Waikoloa Village, Hawaii

Event Website

https://hicss.hawaii.edu/

Start Date

7-1-2025 12:00 AM

End Date

10-1-2025 12:00 AM

Description

In predictive analytics domains, such as healthcare, marketing and finance, data exhibits inherent segmentation, like patient, customer and market segments. Powerful global models, like XGBoost or Catboost, offer high predictive qualities, yet ignore modeling clusters explicitly and are limited by low interpretability. Cluster-then-predict (CTP) models have been proposed to offer more actionable insights. These hybrid models first segment data and then train cluster-specific linear models, combining the capacity to model complex relationships with model transparency. Previous CTP approaches rely on decision trees for segmentation, neglecting alternative methods. This study proposes six CTP models and benchmarks them against five global models. Our results show that k-means CTP ranks fourth out of eleven models in 20 benchmark datasets. While CTP models with DTs rank fifth best, they are substantially simpler to interpret. Consequently, we establish a variety of cluster-then-predict models and call for their consideration when faced with heterogeneous datasets.

Share

COinS
 
Jan 7th, 12:00 AM Jan 10th, 12:00 AM

Benchmarking Cluster-Then-Predict Models to Challenge Prevailing Global Machine Learning Models

Hilton Waikoloa Village, Hawaii

In predictive analytics domains, such as healthcare, marketing and finance, data exhibits inherent segmentation, like patient, customer and market segments. Powerful global models, like XGBoost or Catboost, offer high predictive qualities, yet ignore modeling clusters explicitly and are limited by low interpretability. Cluster-then-predict (CTP) models have been proposed to offer more actionable insights. These hybrid models first segment data and then train cluster-specific linear models, combining the capacity to model complex relationships with model transparency. Previous CTP approaches rely on decision trees for segmentation, neglecting alternative methods. This study proposes six CTP models and benchmarks them against five global models. Our results show that k-means CTP ranks fourth out of eleven models in 20 benchmark datasets. While CTP models with DTs rank fifth best, they are substantially simpler to interpret. Consequently, we establish a variety of cluster-then-predict models and call for their consideration when faced with heterogeneous datasets.

https://aisel.aisnet.org/hicss-58/da/ai_model_evaluation/6