ITEM DISCRIMINATION AND RELIABILITY OF HUMAN-DESIGNED VERSUS AI-GENERATED ACHIEVEMENT TESTS: A CLASSICAL TEST THEORY COMPARISON
Keywords:
artificial intelligence; automated item generation; item discrimination; test reliability; Classical Test Theory; KR-20; graduate assessmentAbstract
The rapid adoption of large language models for educational content creation has outpaced the empirical evidence needed to judge whether AI-generated multiple-choice items meet established psychometric standards, particularly at the graduate level and within under-represented higher-education systems. Using Classical Test Theory, this study addressed a single, focused research question: do human-designed and Claude AI-generated multiple-choice achievement tests, matched for content and item difficulty, differ in item discrimination and internal-consistency reliability? Two 30-item parallel forms were constructed from an identical Table of Specifications aligned with Bloom’s Revised Taxonomy and administered, in a within-subjects design, to 200 graduate students at Kohat University of Science and Technology, Pakistan. Item discrimination was estimated using the upper–lower index and corrected point-biserial correlation; internal consistency was estimated using KR-20. Item difficulty was compared first, as a precondition check, before testing the primary hypothesis. The two forms did not differ significantly in mean item difficulty (human M = .555, AI M = .544; t (29) = 1.324, p = .196), confirming a fair basis for comparison. The AI-generated form showed significantly higher mean discrimination (M = .531 vs. .476; t (29) = −2.384, p = .024) and significantly higher internal-consistency reliability (KR-20 = .850 vs. .806; 95% CI for the difference .015, .076). Under matched content and difficulty conditions, Claude AI-generated items discriminated more effectively between higher- and lower-performing students and produced a more internally consistent achievement test than human-designed items. These findings offer localized, quantitative evidence that generative AI can meet or exceed core Classical Test Theory benchmarks for graduate-level achievement testing, supporting cautious, expert-supervised integration of AI-assisted item development in resource-constrained higher-education contexts.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Abdul Azeem, Prof. Dr. Muhammad Naseer Ud Din, Dr. Abdul Wahab, Shakeel Nawaz

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
This is an open access article under the term of Creative Commons Attribution-NonCommercial 4.0 International license (CC BY-NC 4.0). This license permits the users to use, reproduce, disseminate, or display the article in any medium provided that the authors are the original creators and that the reuse is restricted to non-commercial purposes, i.e., is attributed to research or educational use and the work is also properly cited.