Paper Type
ERF
Abstract
Evaluating technical explanations is an ill-structured task, requiring assigning ratings and justifying judgments without a single ground truth. In multi-agent settings, alignment is typically assessed through agreement, yet similar ratings may arise from different justificatory expressions, limiting insight into how judgments are formed. We examine evaluator alignment using two analytical lenses: convergence (similarity in meaning), measured using cosine similarity, and diagnosticity (structure of information in justifications), measured using entropy difference. We apply this approach in an exploratory study comparing two human cybersecurity experts and two LLM evaluators evaluating 71 short-text explanations under a shared rubric. We find that LLM evaluators exhibit higher convergence and more stable diagnosticity, whereas human evaluators show greater variability, including when assigning the same rating. These findings suggest that agreement does not imply shared justificatory reasoning and highlight the value of analyzing justification texts in human–LLM evaluation systems.
Paper Number
1778
Recommended Citation
Cotoranu, Andreea, "Beyond Agreement: Structural Alignment in Hybrid Human–LLM Evaluation Systems" (2026). AMCIS 2026 Proceedings. 29.
https://aisel.aisnet.org/amcis2026/conftheme/conftheme/29
Beyond Agreement: Structural Alignment in Hybrid Human–LLM Evaluation Systems
Evaluating technical explanations is an ill-structured task, requiring assigning ratings and justifying judgments without a single ground truth. In multi-agent settings, alignment is typically assessed through agreement, yet similar ratings may arise from different justificatory expressions, limiting insight into how judgments are formed. We examine evaluator alignment using two analytical lenses: convergence (similarity in meaning), measured using cosine similarity, and diagnosticity (structure of information in justifications), measured using entropy difference. We apply this approach in an exploratory study comparing two human cybersecurity experts and two LLM evaluators evaluating 71 short-text explanations under a shared rubric. We find that LLM evaluators exhibit higher convergence and more stable diagnosticity, whereas human evaluators show greater variability, including when assigning the same rating. These findings suggest that agreement does not imply shared justificatory reasoning and highlight the value of analyzing justification texts in human–LLM evaluation systems.
When commenting on articles, please be friendly, welcoming, respectful and abide by the AIS eLibrary Discussion Thread Code of Conduct posted here.
Comments
NEXTTRANS