From Investigating the Alignment of A Priori Item Characteristics Based on the CTT and Four-Parameter Logistic (4-PL) IRT Models to Further Exploring the Comparability of the Two Models

Closed

Agus Santoso, Heri Retnawati, Timbul Pardede, Ezi Apino, Ibnu Rafi, Munaya Nikma Rosyada, Gulzhaina K. Kassymova, Xu Wenxin

2024 Practical Assessment, Research and Evaluation Vol. 29 Issue 14 Article Cited by 1 SDG 4SDG 12SDG 8SDG 17 Quartile

Abstract

The test blueprint is important in test development, where it guides the test item writer in creating test items according to the desired objectives and specifications or characteristics (so-called a priori item characteristics), such as the level of item difficulty in the category and the distribution of items based on their difficulty level. Given that the difficulty level of the test items (easy, medium, or hard) created is influenced by the perceptions, knowledge, and experience of the item writer, item analysis based on empirical data using a specific measurement framework needs to be conducted, in addition to evaluation based on expert judgment, to ensure that the test items and the test itself have appropriate characteristics. The present study investigated the extent to which the a priori characteristics (i.e., item difficulty) of the items of the Business English test taken by 4,836 Universitas Terbuka (UT) students aligned with their characteristics when estimated under classical test theory (CTT) and four-parameter logistic (4-PL) IRT models based on empirical data. In light of the two measurement models used, CTT and 4-PL, we extended this study to exploring the comparability of the two models based on the yielded item difficulty and discrimination estimates and the relationship between pseudo-guessing and carelessness parameters. Our study suggested insufficient support for asserting that the characteristics of the items used in the Business English test align with the characteristics expected by the test developers. The exploration of the comparability of the CTT and 4-PL models demonstrated that while the two models were comparable in terms of the item difficulty estimates yielded, they were not comparable for the item discrimination estimates. Our study also did not find a linear association of the pseudo-guessing and carelessness parameters estimated under the 4-PL model. Further findings of our study and their implications, especially on test development practices, are discussed. Copyright is retained by the first or sole author, who grants right of first publication to Practical Assessment, Research & Evaluation. Permission is granted to distribute this article for nonprofit, educational purposes if it is copied in its entirety and the journal is credited. PARE has the right to authorize third party reproduction of this article in print, electronic and database forms.

Affiliations

Universitas Terbuka, Indonesia; Universitas Negeri Yogyakarta, Indonesia; Abai Kazakh National Pedagogical University, Kazakhstan

Research at a Glance

Premium content — register to unlock

Research at a Glance

Register to unlock

Topics & SDG Alignment

Premium content — register to unlock

Topics & SDG Alignment

Register to unlock

Collaboration

Premium content — register to unlock

Collaboration

Register to unlock

Author Profile (Selected)

Premium content — register to unlock

Author Profile (Selected)

Register to unlock

References Overview

Premium content — register to unlock

References Overview

Register to unlock

Journal & Source

Premium content — register to unlock

Journal & Source

Register to unlock

Metadata & Integrity

Premium content — register to unlock

Metadata & Integrity

Register to unlock