Abstract

This study investigates machine learning approaches for toxicity prediction using the Tox21 dataset, which contains 11,760 chemical molecules represented as SMILES sequences and 12 binary toxicity prediction tasks. The dataset is characterized by class imbalance and missing labels, making toxicity prediction particularly challenging. Two approaches were evaluated: a single-task transformer-based model (ChemBERTa) and a multi-task multilayer perceptron (MLP) using Morgan fingerprints. The experiments were conducted using stratified 5-fold cross-validation, with ROC-AUC as the primary evaluation metric. The results indicate that the multi-task approach achieved higher predictive performance (ROC-AUC = 0.8623) than the single-task ChemBERTa model (ROC-AUC = 0.8488), while requiring substantially lower computational resources. Additionally, for the evaluated SR-HSE task, partial fine-tuning of ChemBERTa reduced the number of trainable parameters from 44.1M to 14.8M with only a moderate decrease in predictive performance.

Recommended Citation

Frączkowska, A. & Jastrzebska, A.(2026). Single-Task vs Multi-Task Learning for Toxicity Prediction: A Study on Tox21. In M. Valenta, B. Mannová, R. Pergl, A. Przybylek, M. Lang, H. Linger, C. Schneider, N. Iivari, & E. Insfran (Eds.), Making ISD Sustainable: Reloaded with AI and Automation (ISD2026 Proceedings). Prague, Czech Republic: Czech Technical University in Prague. ISBN: 978-80-01-07585-2. https://doi.org/10.62036/ISD.2026.98

Paper Type

Poster

DOI

10.62036/ISD.2026.98

Share

COinS
 

Single-Task vs Multi-Task Learning for Toxicity Prediction: A Study on Tox21

This study investigates machine learning approaches for toxicity prediction using the Tox21 dataset, which contains 11,760 chemical molecules represented as SMILES sequences and 12 binary toxicity prediction tasks. The dataset is characterized by class imbalance and missing labels, making toxicity prediction particularly challenging. Two approaches were evaluated: a single-task transformer-based model (ChemBERTa) and a multi-task multilayer perceptron (MLP) using Morgan fingerprints. The experiments were conducted using stratified 5-fold cross-validation, with ROC-AUC as the primary evaluation metric. The results indicate that the multi-task approach achieved higher predictive performance (ROC-AUC = 0.8623) than the single-task ChemBERTa model (ROC-AUC = 0.8488), while requiring substantially lower computational resources. Additionally, for the evaluated SR-HSE task, partial fine-tuning of ChemBERTa reduced the number of trainable parameters from 44.1M to 14.8M with only a moderate decrease in predictive performance.