Location
Hilton Waikoloa Village, Hawaii
Event Website
https://hicss.hawaii.edu/
Start Date
7-1-2025 12:00 AM
End Date
10-1-2025 12:00 AM
Description
BERT-based models have become the mainstream in sentiment classification approaches. However, due to the divergence of the text domains, each domain requires a specific fine-tuned model which is often impractical for scaling. Additionally, the large model's size requires heavy computation resources. In our work, we propose a framework which could address such issues by combining the domain adaptation task with a lightweight model distillation. From each trained model of a specific domain, a merged model is created by fusing all models without the need to finetune on a combined dataset. Consequentially, the resulting model is distilled into a smaller model to lower the required computation. We test our framework on semantic classification with Vietnamese datasets with a pre-trained BERT-based architecture. The results highlight that our merged model achieves the highest average accuracy overall substantially while the distilled model maintains a competitive performance with a 50\%{} reduction in inference time.
Recommended Citation
Tran, Ngoc Minh; Le, Bang Giang; and Ta, Viet Cuong, "MergeKD: An Empirical Framework for Combining Knowledge Distillation with Model Fusion Using BERT Model" (2025). Hawaii International Conference on System Sciences 2025 (HICSS-58). 6.
https://aisel.aisnet.org/hicss-58/st/edge_computing/6
MergeKD: An Empirical Framework for Combining Knowledge Distillation with Model Fusion Using BERT Model
Hilton Waikoloa Village, Hawaii
BERT-based models have become the mainstream in sentiment classification approaches. However, due to the divergence of the text domains, each domain requires a specific fine-tuned model which is often impractical for scaling. Additionally, the large model's size requires heavy computation resources. In our work, we propose a framework which could address such issues by combining the domain adaptation task with a lightweight model distillation. From each trained model of a specific domain, a merged model is created by fusing all models without the need to finetune on a combined dataset. Consequentially, the resulting model is distilled into a smaller model to lower the required computation. We test our framework on semantic classification with Vietnamese datasets with a pre-trained BERT-based architecture. The results highlight that our merged model achieves the highest average accuracy overall substantially while the distilled model maintains a competitive performance with a 50\%{} reduction in inference time.
https://aisel.aisnet.org/hicss-58/st/edge_computing/6