Abstract

Large language models (LLMs) are increasingly used to detect and evaluate sentiment in unstructured text. Configuring an LLM annotation pipeline requires researchers to make numerous design choices, raising the question of how sensitive annotation outputs are to these researcher degrees of freedom. This study investigates the reliability and validity of LLM annotations by systematically varying five pipeline parameters across four open-weight models, 150 multilingual hotel reviews, and six service quality attributes, producing over 61,000 annotation data points. Results show that annotations are highly reliable and largely insensitive to pipeline configuration, with internal consistency well above acceptability thresholds (Krippendorff’s α =0.839–0.944). Criterion validity is also high, with LLM composite scores correlating with external ratings at ρ = 0.816–0.874, significantly outperforming both VADER and multilingual BERT baselines.

Recommended Citation

Smolinski, P., Bonin, A.L., Staegemann, D., Pohl, M. & Winiarski, J.(2026). Reliability and Validity of LLM-Based Text Annotation. In M. Valenta, B. Mannová, R. Pergl, A. Przybylek, M. Lang, H. Linger, C. Schneider, N. Iivari, & E. Insfran (Eds.), Making ISD Sustainable: Reloaded with AI and Automation (ISD2026 Proceedings). Prague, Czech Republic: Czech Technical University in Prague. ISBN: 978-80-01-07585-2. https://doi.org/10.62036/ISD.2026.199

Paper Type

Poster

DOI

10.62036/ISD.2026.199

Share

COinS
 

Reliability and Validity of LLM-Based Text Annotation

Large language models (LLMs) are increasingly used to detect and evaluate sentiment in unstructured text. Configuring an LLM annotation pipeline requires researchers to make numerous design choices, raising the question of how sensitive annotation outputs are to these researcher degrees of freedom. This study investigates the reliability and validity of LLM annotations by systematically varying five pipeline parameters across four open-weight models, 150 multilingual hotel reviews, and six service quality attributes, producing over 61,000 annotation data points. Results show that annotations are highly reliable and largely insensitive to pipeline configuration, with internal consistency well above acceptability thresholds (Krippendorff’s α =0.839–0.944). Criterion validity is also high, with LLM composite scores correlating with external ratings at ρ = 0.816–0.874, significantly outperforming both VADER and multilingual BERT baselines.