Paper Type

Short

Paper Number

PACIS2026-1490

Description

This paper introduces a Knowledge‑State Generative Agent framework for evaluating the quality of pre‑assessment questions. The framework employs large language model (LLM)–based agents prompted to adopt a teacher persona to simulate the responses of students with and without mastery of targeted knowledge components. A preliminary empirical study using archival data from 424 students enrolled in an Information Systems Management course indicates that the proposed approach yields interpretable metrics under Classical Test Theory. Results further show that agents instantiated with the relevant mastered knowledge components exhibit systematically higher performance than agents lacking such mastery. In addition, the study suggests that teacher-persona prompting may mitigate the high-accuracy bias observed in LLMs, whereby models tend to produce correct answers even when prompted to disregard latent knowledge acquired during training. Future work will extend the framework beyond multiple-choice questions and evaluate its generalizability across interdisciplinary domains.

Comments

01-AIML

Share

COinS
 
Jul 5th, 12:00 AM

Knowledge-State Generative Agents for Pre Assessment Question Evaluation

This paper introduces a Knowledge‑State Generative Agent framework for evaluating the quality of pre‑assessment questions. The framework employs large language model (LLM)–based agents prompted to adopt a teacher persona to simulate the responses of students with and without mastery of targeted knowledge components. A preliminary empirical study using archival data from 424 students enrolled in an Information Systems Management course indicates that the proposed approach yields interpretable metrics under Classical Test Theory. Results further show that agents instantiated with the relevant mastered knowledge components exhibit systematically higher performance than agents lacking such mastery. In addition, the study suggests that teacher-persona prompting may mitigate the high-accuracy bias observed in LLMs, whereby models tend to produce correct answers even when prompted to disregard latent knowledge acquired during training. Future work will extend the framework beyond multiple-choice questions and evaluate its generalizability across interdisciplinary domains.