Abstract

Graph kernels are a powerful tool for analyzing graph datasets. Most existing approaches focus solely on optimizing performance metrics rather than model explanation quality. In this paper, we introduce the Generalized Graphlet kernel, an extension of the classical graphlet kernel. Unlike its predecessor, it optimizes both classification accuracy and the quality of model explanations. The resulting Pareto-optimal subset of graphlets becomes an explanation for the model, making the model itself a source of knowledge, not just the input dataset. We evaluate our method on 12 benchmark datasets for biochemical graph classification using AdaBoost, Random Forest, and Support Vector Machine classifiers. Our results show that our method consistently improves explanation quality and classification accuracy compared to the standard graphlet kernel, particularly for SVC and AdaBoost. The Pareto-optimal graphlets provide interpretable, domain-relevant patterns, demonstrating that our method creates models that are both accurate and explainable.

Recommended Citation

Kuhar, Y. & Čibej, U.(2026). Optimizing explainability of graphlet kernels. In M. Valenta, B. Mannová, R. Pergl, A. Przybylek, M. Lang, H. Linger, C. Schneider, N. Iivari, & E. Insfran (Eds.), Making ISD Sustainable: Reloaded with AI and Automation (ISD2026 Proceedings). Prague, Czech Republic: Czech Technical University in Prague. ISBN: 978-80-01-07585-2. https://doi.org/10.62036/ISD.2026.86

Paper Type

Full Paper

DOI

10.62036/ISD.2026.86

Share

COinS
 

Optimizing explainability of graphlet kernels

Graph kernels are a powerful tool for analyzing graph datasets. Most existing approaches focus solely on optimizing performance metrics rather than model explanation quality. In this paper, we introduce the Generalized Graphlet kernel, an extension of the classical graphlet kernel. Unlike its predecessor, it optimizes both classification accuracy and the quality of model explanations. The resulting Pareto-optimal subset of graphlets becomes an explanation for the model, making the model itself a source of knowledge, not just the input dataset. We evaluate our method on 12 benchmark datasets for biochemical graph classification using AdaBoost, Random Forest, and Support Vector Machine classifiers. Our results show that our method consistently improves explanation quality and classification accuracy compared to the standard graphlet kernel, particularly for SVC and AdaBoost. The Pareto-optimal graphlets provide interpretable, domain-relevant patterns, demonstrating that our method creates models that are both accurate and explainable.