Location

Hilton Waikoloa Village, Hawaii

Event Website

https://hicss.hawaii.edu/

Start Date

7-1-2025 12:00 AM

End Date

10-1-2025 12:00 AM

Description

BERT-based models have become the mainstream in sentiment classification approaches. However, due to the divergence of the text domains, each domain requires a specific fine-tuned model which is often impractical for scaling. Additionally, the large model's size requires heavy computation resources. In our work, we propose a framework which could address such issues by combining the domain adaptation task with a lightweight model distillation. From each trained model of a specific domain, a merged model is created by fusing all models without the need to finetune on a combined dataset. Consequentially, the resulting model is distilled into a smaller model to lower the required computation. We test our framework on semantic classification with Vietnamese datasets with a pre-trained BERT-based architecture. The results highlight that our merged model achieves the highest average accuracy overall substantially while the distilled model maintains a competitive performance with a 50\%{} reduction in inference time.

Share

COinS
 
Jan 7th, 12:00 AM Jan 10th, 12:00 AM

MergeKD: An Empirical Framework for Combining Knowledge Distillation with Model Fusion Using BERT Model

Hilton Waikoloa Village, Hawaii

BERT-based models have become the mainstream in sentiment classification approaches. However, due to the divergence of the text domains, each domain requires a specific fine-tuned model which is often impractical for scaling. Additionally, the large model's size requires heavy computation resources. In our work, we propose a framework which could address such issues by combining the domain adaptation task with a lightweight model distillation. From each trained model of a specific domain, a merged model is created by fusing all models without the need to finetune on a combined dataset. Consequentially, the resulting model is distilled into a smaller model to lower the required computation. We test our framework on semantic classification with Vietnamese datasets with a pre-trained BERT-based architecture. The results highlight that our merged model achieves the highest average accuracy overall substantially while the distilled model maintains a competitive performance with a 50\%{} reduction in inference time.

https://aisel.aisnet.org/hicss-58/st/edge_computing/6