ITEM DISCRIMINATION AND RELIABILITY OF HUMAN-DESIGNED VERSUS AI-GENERATED ACHIEVEMENT TESTS: A CLASSICAL TEST THEORY COMPARISON

Authors

  • Abdul Azeem M.Phil. Scholar Institute of Education and Research, Kohat University of Science and Technology
  • Prof. Dr. Muhammad Naseer Ud Din Director, Institute of Education and Research, Kohat University of Science and Technology.
  • Dr. Abdul Wahab Lecturer, Institute of Education and Research, Kohat University of Science and Technology
  • Shakeel Nawaz PhD scholar/Visiting Lecturer, Institute of Education and Research, Kohat University of Science and Technology

Keywords:

artificial intelligence; automated item generation; item discrimination; test reliability; Classical Test Theory; KR-20; graduate assessment

Abstract

The rapid adoption of large language models for educational content creation has outpaced the empirical evidence needed to judge whether AI-generated multiple-choice items meet established psychometric standards, particularly at the graduate level and within under-represented higher-education systems. Using Classical Test Theory, this study addressed a single, focused research question: do human-designed and Claude AI-generated multiple-choice achievement tests, matched for content and item difficulty, differ in item discrimination and internal-consistency reliability? Two 30-item parallel forms were constructed from an identical Table of Specifications aligned with Bloom’s Revised Taxonomy and administered, in a within-subjects design, to 200 graduate students at Kohat University of Science and Technology, Pakistan. Item discrimination was estimated using the upper–lower index and corrected point-biserial correlation; internal consistency was estimated using KR-20. Item difficulty was compared first, as a precondition check, before testing the primary hypothesis. The two forms did not differ significantly in mean item difficulty (human M = .555, AI M = .544; t (29) = 1.324, p = .196), confirming a fair basis for comparison. The AI-generated form showed significantly higher mean discrimination (M = .531 vs. .476; t (29) = −2.384, p = .024) and significantly higher internal-consistency reliability (KR-20 = .850 vs. .806; 95% CI for the difference .015, .076). Under matched content and difficulty conditions, Claude AI-generated items discriminated more effectively between higher- and lower-performing students and produced a more internally consistent achievement test than human-designed items. These findings offer localized, quantitative evidence that generative AI can meet or exceed core Classical Test Theory benchmarks for graduate-level achievement testing, supporting cautious, expert-supervised integration of AI-assisted item development in resource-constrained higher-education contexts.

Downloads

Published

2026-07-20

How to Cite

Abdul Azeem, Prof. Dr. Muhammad Naseer Ud Din, Dr. Abdul Wahab, & Shakeel Nawaz. (2026). ITEM DISCRIMINATION AND RELIABILITY OF HUMAN-DESIGNED VERSUS AI-GENERATED ACHIEVEMENT TESTS: A CLASSICAL TEST THEORY COMPARISON. Review of Crime, Peace and Society, 3(4), 55–60. Retrieved from https://reviewcps.com/index.php/rcps/article/view/114

Similar Articles

1 2 3 4 > >> 

You may also start an advanced similarity search for this article.