Abstract

This paper does not aim to present a clinically ready Parkinson’s disease classifier. Instead, it rigorously examines what actually works in an extreme low-sample setting for bilateral, smoothed sEMG signals, using a dataset of only 9 subjects (5 Parkinson’s disease, 4 healthy controls; 18 full recordings, 122 segments). We audit classical subject-level baselines and multiple state-graph families under leakage-safe protocols, with seed-stability analysis, segment augmentation, paired tests, and a permutation sanity check.

A key finding is that isolated strong graph-model results can be misleading: some frozen graph configurations achieved ROC–AUC of 1.0, yet did not remain stable across seed reruns, and the best-performing graph branch changed across audited subsets. After fixing an evaluation bug, a naive subject-level baseline reached only 0.55 ROC–AUC, whereas a leakage-safe nested sparse baseline (outer and inner LOSO, fold-wise feature selection, logistic regression) achieved 0.778 accuracy, 0.775 balanced accuracy, 0.800 F1, 0.550 MCC, and 0.900 ROC–AUC, using 5.89 features per fold on average. The contribution is methodological: in ultra-low-N settings, evaluation rigor and stability analysis determine what counts as a valid result.

Recommended Citation

Cieślak, D., Florkiewicz, L. & Kaczmarek, M.(2026). What Actually Works in Ultra-Small sEMG Datasets? A Comprehensive Audit of Classical and Graph Models for Subject-Level Parkinson's Detection. In M. Valenta, B. Mannová, R. Pergl, A. Przybylek, M. Lang, H. Linger, C. Schneider, N. Iivari, & E. Insfran (Eds.), Making ISD Sustainable: Reloaded with AI and Automation (ISD2026 Proceedings). Prague, Czech Republic: Czech Technical University in Prague. ISBN: 978-80-01-07585-2. https://doi.org/10.62036/ISD.2026.107

Paper Type

Short Paper

DOI

10.62036/ISD.2026.107

Share

COinS
 

What Actually Works in Ultra-Small sEMG Datasets? A Comprehensive Audit of Classical and Graph Models for Subject-Level Parkinson's Detection

This paper does not aim to present a clinically ready Parkinson’s disease classifier. Instead, it rigorously examines what actually works in an extreme low-sample setting for bilateral, smoothed sEMG signals, using a dataset of only 9 subjects (5 Parkinson’s disease, 4 healthy controls; 18 full recordings, 122 segments). We audit classical subject-level baselines and multiple state-graph families under leakage-safe protocols, with seed-stability analysis, segment augmentation, paired tests, and a permutation sanity check.

A key finding is that isolated strong graph-model results can be misleading: some frozen graph configurations achieved ROC–AUC of 1.0, yet did not remain stable across seed reruns, and the best-performing graph branch changed across audited subsets. After fixing an evaluation bug, a naive subject-level baseline reached only 0.55 ROC–AUC, whereas a leakage-safe nested sparse baseline (outer and inner LOSO, fold-wise feature selection, logistic regression) achieved 0.778 accuracy, 0.775 balanced accuracy, 0.800 F1, 0.550 MCC, and 0.900 ROC–AUC, using 5.89 features per fold on average. The contribution is methodological: in ultra-low-N settings, evaluation rigor and stability analysis determine what counts as a valid result.