Paper Type
Short
Paper Number
PACIS2026-1901
Description
Online forums provide rich insights into real-world user experiences with generative artificial intelligence (GenAI), yet the highly unstructured and noisy nature of social media text poses significant challenges for automated analysis. Our study proposes an automated annotation framework that integrates Snorkel-based weak supervision with semantic embedding models to identify valuable information from high-noise discussions. Using data collected from Taiwan’s PTT forum, we construct semantic anchors based on the BGE embedding model to capture contextual meanings beyond traditional keyword rules. These anchors generate weak labels that are integrated through the Snorkel framework and subsequently used to train a LightGBM classifier. Experimental results show that the proposed approach achieves strong classification performance, reaching 87% accuracy, Macro F1 of 0.87, and AUC of 0.922. The findings demonstrate that combining semantic embeddings with weak supervision can effectively transform noisy social media data into actionable knowledge while significantly reducing the need for large-scale manual annotation.
Recommended Citation
Lai, Chiayu; Yang, Yu-Chen; Chen, Deng-Neng; and Lin, Chih-Ting, "From Useless Noise to Useful Information: Automated Social Media Annotation in High-Noise Environments" (2026). PACIS 2026 Proceedings. 15.
https://aisel.aisnet.org/pacis2026/general_topic/general_topic/15
From Useless Noise to Useful Information: Automated Social Media Annotation in High-Noise Environments
Online forums provide rich insights into real-world user experiences with generative artificial intelligence (GenAI), yet the highly unstructured and noisy nature of social media text poses significant challenges for automated analysis. Our study proposes an automated annotation framework that integrates Snorkel-based weak supervision with semantic embedding models to identify valuable information from high-noise discussions. Using data collected from Taiwan’s PTT forum, we construct semantic anchors based on the BGE embedding model to capture contextual meanings beyond traditional keyword rules. These anchors generate weak labels that are integrated through the Snorkel framework and subsequently used to train a LightGBM classifier. Experimental results show that the proposed approach achieves strong classification performance, reaching 87% accuracy, Macro F1 of 0.87, and AUC of 0.922. The findings demonstrate that combining semantic embeddings with weak supervision can effectively transform noisy social media data into actionable knowledge while significantly reducing the need for large-scale manual annotation.
Comments
17-General