Paper Type
Short
Paper Number
PACIS2026-1490
Description
This paper introduces a Knowledge‑State Generative Agent framework for evaluating the quality of pre‑assessment questions. The framework employs large language model (LLM)–based agents prompted to adopt a teacher persona to simulate the responses of students with and without mastery of targeted knowledge components. A preliminary empirical study using archival data from 424 students enrolled in an Information Systems Management course indicates that the proposed approach yields interpretable metrics under Classical Test Theory. Results further show that agents instantiated with the relevant mastered knowledge components exhibit systematically higher performance than agents lacking such mastery. In addition, the study suggests that teacher-persona prompting may mitigate the high-accuracy bias observed in LLMs, whereby models tend to produce correct answers even when prompted to disregard latent knowledge acquired during training. Future work will extend the framework beyond multiple-choice questions and evaluate its generalizability across interdisciplinary domains.
Recommended Citation
Ke, Ping Fan; Lau, Yi Meng; and Lo, Siaw Ling, "Knowledge-State Generative Agents for Pre Assessment Question Evaluation" (2026). PACIS 2026 Proceedings. 7.
https://aisel.aisnet.org/pacis2026/ai_ml/ai_ml/7
Knowledge-State Generative Agents for Pre Assessment Question Evaluation
This paper introduces a Knowledge‑State Generative Agent framework for evaluating the quality of pre‑assessment questions. The framework employs large language model (LLM)–based agents prompted to adopt a teacher persona to simulate the responses of students with and without mastery of targeted knowledge components. A preliminary empirical study using archival data from 424 students enrolled in an Information Systems Management course indicates that the proposed approach yields interpretable metrics under Classical Test Theory. Results further show that agents instantiated with the relevant mastered knowledge components exhibit systematically higher performance than agents lacking such mastery. In addition, the study suggests that teacher-persona prompting may mitigate the high-accuracy bias observed in LLMs, whereby models tend to produce correct answers even when prompted to disregard latent knowledge acquired during training. Future work will extend the framework beyond multiple-choice questions and evaluate its generalizability across interdisciplinary domains.
Comments
01-AIML