The Experts below are selected from a list of 3048 Experts worldwide ranked by ideXlab platform
Huahua Chang - One of the best experts on this subject based on the ideXlab platform.
-
item selection criteria with practical constraints in cognitive diagnostic Computerized Adaptive Testing
Educational and Psychological Measurement, 2019Co-Authors: Chuanju Lin, Huahua ChangAbstract:For item selection in cognitive diagnostic Computerized Adaptive Testing (CD-CAT), ideally, a single item selection index should be created to simultaneously regulate precision, exposure status, and attribute balancing. For this purpose, in this study, we first proposed an attribute-balanced item selection criterion, namely, the standardized weighted deviation global discrimination index (SWDGDI), and subsequently formulated the constrained progressive index (CP_SWDGDI) by casting the SWDGDI in a progressive algorithm. A simulation study revealed that the SWDGDI method was effective in balancing attribute coverage and the CP_SWDGDI method was able to simultaneously balance attribute coverage and item pool usage while maintaining acceptable estimation precision. This research also demonstrates the advantage of a relatively low number of attributes in CD-CAT applications.
-
statistical foundations for Computerized Adaptive Testing with response revision
Psychometrika, 2019Co-Authors: Shiyu Wang, Georgios Fellouris, Huahua ChangAbstract:The compatibility of Computerized Adaptive Testing (CAT) with response revision has been a topic of debate in psychometrics for many years. The problem is to provide test takers opportunities to change their answers during the test, while discouraging deceptive strategies from their side and preserving the statistical efficiency of the traditional CAT. The estimating approach proposed in Wang et al. (Stat Sin 27(4):1987-2010, 2017), based on the nominal response model, allows test takers to provide more than one answer to each item during the test, which they all contribute to the interim and final ability estimation. This approach is here reformulated, extended to incorporate a larger class of polytomous and dichotomous item response theory models, and investigated with simulation studies under different test-taking strategies.
-
sequential detection of compromised items using response times in Computerized Adaptive Testing
Psychometrika, 2018Co-Authors: Edison M Choe, Jinming Zhang, Huahua ChangAbstract:Item compromise persists in undermining the integrity of Testing, even secure administrations of Computerized Adaptive Testing (CAT) with sophisticated item exposure controls. In ongoing efforts to tackle this perennial security issue in CAT, a couple of recent studies investigated sequential procedures for detecting compromised items, in which a significant increase in the proportion of correct responses for each item in the pool is monitored in real time using moving averages. In addition to actual responses, response times are valuable information with tremendous potential to reveal items that may have been leaked. Specifically, examinees that have preknowledge of an item would likely respond more quickly to it than those who do not. Therefore, the current study proposes several augmented methods for the detection of compromised items, all involving simultaneous monitoring of changes in both the proportion correct and average response time for every item using various moving average strategies. Simulation results with an operational item pool indicate that, compared to the analysis of responses alone, utilizing response times can afford marked improvements in detection power with fewer false positives.
-
dual objective item selection criteria in cognitive diagnostic Computerized Adaptive Testing
Journal of Educational Measurement, 2017Co-Authors: Hyeon Ah Kang, Susu Zhang, Huahua ChangAbstract:The development of cognitive diagnostic-Computerized Adaptive Testing (CD-CAT) has provided a new perspective for gaining information about examinees' mastery on a set of cognitive attributes. This study proposes a new item selection method within the framework of dual-objective CD-CAT that simultaneously addresses examinees' attribute mastery status and overall test performance. The new procedure is based on the Jensen-Shannon (JS) divergence, a symmetrized version of the Kullback-Leibler divergence. We show that the JS divergence resolves the noncomparability problem of the dual information index and has close relationships with Shannon entropy, mutual information, and Fisher information. The performance of the JS divergence is evaluated in simulation studies in comparison with the methods available in the literature. Results suggest that the JS divergence achieves parallel or more precise recovery of latent trait variables compared to the existing methods and maintains practical advantages in computation and item pool usage.
-
developing new online calibration methods for multidimensional Computerized Adaptive Testing
British Journal of Mathematical and Statistical Psychology, 2017Co-Authors: Ping Chen, Tao Xin, Chun Wang, Huahua ChangAbstract:Multidimensional Computerized Adaptive Testing (MCAT) has received increasing attention over the past few years in educational measurement. Like all other formats of CAT, item replenishment is an essential part of MCAT for its item bank maintenance and management, which governs retiring overexposed or obsolete items over time and replacing them with new ones. Moreover, calibration precision of the new items will directly affect the estimation accuracy of examinees' ability vectors. In unidimensional CAT (UCAT) and cognitive diagnostic CAT, online calibration techniques have been developed to effectively calibrate new items. However, there has been very little discussion of online calibration in MCAT in the literature. Thus, this paper proposes new online calibration methods for MCAT based upon some popular methods used in UCAT. Three representative methods, Method A, the 'one EM cycle' method and the 'multiple EM cycles' method, are generalized to MCAT. Three simulation studies were conducted to compare the three new methods by manipulating three factors (test length, item bank design, and level of correlation between coordinate dimensions). The results showed that all the new methods were able to recover the item parameters accurately, and the Adaptive online calibration designs showed some improvements compared to the random design under most conditions.
Chun Wang - One of the best experts on this subject based on the ideXlab platform.
-
variable length stopping rules for multidimensional Computerized Adaptive Testing
Psychometrika, 2019Co-Authors: Chun Wang, David J Weiss, Zhuoran ShangAbstract:In Computerized Adaptive Testing (CAT), a variable-length stopping rule refers to ending item administration after a pre-specified measurement precision standard has been satisfied. The goal is to provide equal measurement precision for all examinees regardless of their true latent trait level. Several stopping rules have been proposed in unidimensional CAT, such as the minimum information rule or the maximum standard error rule. These rules have also been extended to multidimensional CAT and cognitive diagnostic CAT, and they all share the same idea of monitoring measurement error. Recently, Babcock and Weiss (J Comput Adapt Test 2012. https://doi.org/10.7333/1212-0101001) proposed an “absolute change in theta” (CT) rule, which is useful when an item bank is exhaustive of good items for one or more ranges of the trait continuum. Choi, Grady and Dodd (Educ Psychol Meas 70:1–17, 2010) also argued that a CAT should stop when the standard error does not change, implying that the item bank is likely exhausted. Although these stopping rules have been evaluated and compared in different simulation studies, the relationships among the various rules remain unclear, and therefore there lacks a clear guideline regarding when to use which rule. This paper presents analytic results to show the connections among various stopping rules within both unidimensional and multidimensional CAT. In particular, it is argued that the CT-rule alone can be unstable and it can end the test prematurely. However, the CT-rule can be a useful secondary rule to monitor the point of diminished returns. To further provide empirical evidence, three simulation studies are reported using both the 2PL model and the multidimensional graded response model.
-
a continuous a stratification index for item exposure control in Computerized Adaptive Testing
Applied Psychological Measurement, 2018Co-Authors: Alan Huebner, Chun Wang, Bridget Daly, Colleen PinkelmanAbstract:The method of a-stratification aims to reduce item overexposure in Computerized Adaptive Testing, as items that are administered at very high rates may threaten the validity of test scores. In existing methods of a-stratification, the item bank is partitioned into a fixed number of nonoverlapping strata according to the items’a, or discrimination, parameters. This article introduces a continuous a-stratification index which incorporates exposure control into the item selection index itself and thus eliminates the need for fixed discrete strata. The new continuous a-stratification index is compared with existing stratification methods via simulation studies in terms of ability estimation bias, mean squared error, and control of item exposure rates.
-
developing new online calibration methods for multidimensional Computerized Adaptive Testing
British Journal of Mathematical and Statistical Psychology, 2017Co-Authors: Ping Chen, Tao Xin, Chun Wang, Huahua ChangAbstract:Multidimensional Computerized Adaptive Testing (MCAT) has received increasing attention over the past few years in educational measurement. Like all other formats of CAT, item replenishment is an essential part of MCAT for its item bank maintenance and management, which governs retiring overexposed or obsolete items over time and replacing them with new ones. Moreover, calibration precision of the new items will directly affect the estimation accuracy of examinees' ability vectors. In unidimensional CAT (UCAT) and cognitive diagnostic CAT, online calibration techniques have been developed to effectively calibrate new items. However, there has been very little discussion of online calibration in MCAT in the literature. Thus, this paper proposes new online calibration methods for MCAT based upon some popular methods used in UCAT. Three representative methods, Method A, the 'one EM cycle' method and the 'multiple EM cycles' method, are generalized to MCAT. Three simulation studies were conducted to compare the three new methods by manipulating three factors (test length, item bank design, and level of correlation between coordinate dimensions). The results showed that all the new methods were able to recover the item parameters accurately, and the Adaptive online calibration designs showed some improvements compared to the random design under most conditions.
-
on initial item selection in cognitive diagnostic Computerized Adaptive Testing
British Journal of Mathematical and Statistical Psychology, 2016Co-Authors: Chun Wang, Zhuoran ShangAbstract:There has recently been much interest in Computerized Adaptive Testing (CAT) for cognitive diagnosis. While there exist various item selection criteria and different asymptotically optimal designs, these are mostly constructed based on the asymptotic theory assuming the test length goes to infinity. In practice, with limited test lengths, the desired asymptotic optimality may not always apply, and there are few studies in the literature concerning the optimal design of finite items. Related questions, such as how many items we need in order to be able to identify the attribute pattern of an examinee and what types of initial items provide the optimal classification results, are still open. This paper aims to answer these questions by providing non-asymptotic theory of the optimal selection of initial items in cognitive diagnostic CAT. In particular, for the optimal design, we provide necessary and sufficient conditions for the Q-matrix structure of the initial items. The theoretical development is suitable for a general family of cognitive diagnostic models. The results not only provide a guideline for the design of optimal item selection procedures, but also may be applied to guide item bank construction.
-
a new online calibration method for multidimensional Computerized Adaptive Testing
Psychometrika, 2016Co-Authors: Ping Chen, Chun WangAbstract:Multidimensional-Method A (M-Method A) has been proposed as an efficient and effective online calibration method for multidimensional Computerized Adaptive Testing (MCAT) (Chen & Xin, Paper presented at the 78th Meeting of the Psychometric Society, Arnhem, The Netherlands, 2013). However, a key assumption of M-Method A is that it treats person parameter estimates as their true values, thus this method might yield erroneous item calibration when person parameter estimates contain non-ignorable measurement errors. To improve the performance of M-Method A, this paper proposes a new MCAT online calibration method, namely, the full functional MLE-M-Method A (FFMLE-M-Method A). This new method combines the full functional MLE (Jones & Jin in Psychometrika 59:59-75, 1994; Stefanski & Carroll in Annals of Statistics 13:1335-1351, 1985) with the original M-Method A in an effort to correct for the estimation error of ability vector that might otherwise adversely affect the precision of item calibration. Two correction schemes are also proposed when implementing the new method. A simulation study was conducted to show that the new method generated more accurate item parameter estimation than the original M-Method A in almost all conditions.
Tao Xin - One of the best experts on this subject based on the ideXlab platform.
-
the block item pocket method for reviewable multidimensional Computerized Adaptive Testing
Applied Psychological Measurement, 2021Co-Authors: Zhe Lin, Ping Chen, Tao XinAbstract:Most Computerized Adaptive Testing (CAT) programs do not allow item review due to a decrease in estimation precision and aberrant manipulation strategies. In this article, a block item pocket (BIP)...
-
Binary Restrictive Threshold Method for Item Exposure Control in Cognitive Diagnostic Computerized Adaptive Testing
'Frontiers Media SA', 2021Co-Authors: Tao Xin, Xiaojian Sun, Yizhu Gao, Naiqing SongAbstract:Although classification accuracy is a critical issue in cognitive diagnostic Computerized Adaptive Testing, attention has increasingly shifted to item exposure control to ensure test security. In this study, we developed the binary restrictive threshold (BRT) method to balance measurement accuracy and item exposure. In addition, a simulation study was conducted to evaluate its performance. The results indicated that the BRT method performed better than the restrictive progressive (RP) and stratified dynamic binary searching (SDBS) approaches but worse than the restrictive threshold (RT) method in terms of classification accuracy. With respect to item exposure control, the BRT method exhibited noticeably stronger performance compared with the RT method, even though its performance was not as high as that of the RP and SDBS methods
-
attribute discrimination index based method to balance attribute coverage for short length cognitive diagnostic Computerized Adaptive Testing
Frontiers in Psychology, 2020Co-Authors: Yutong Wang, Xiaojian Sun, Weifeng Chong, Tao XinAbstract:We propose a new method that balances attribute coverage for short-length cognitive diagnostic Computerized Adaptive Testing (CD-CAT). The new method uses the attribute discrimination index (ADI-based method) instead of the number of items that measure each attribute [modified global discrimination index (MGDI)-based method] to balance the attribute coverage. Therefore, the information that each attribute provides can be captured. The purpose of the simulation study was to evaluate the performance of the new method, and the results showed the following: (a) Compared with uncontrolled attribute-balance coverage method, the new method produced a higher mastery pattern correct classification rate (PCCR) and attribute correct classification rate (ACCR) with both the posterior-weighted Kullback-Leibler (PWKL) and the modified PWKL (MPWKL) item selection method. (b) Equalization of ACCR (E-ACCR) based on the ADI-based method leads to better results, followed by the MGDI-based method. The uncontrolled method leads to the worst results regardless of item selection methods. (c) Both the ADI-based and MGDI-based methods produced acceptable examinee qualification rates, regardless of item selection methods, although they were relatively low for the uncontrolled condition.
-
developing new online calibration methods for multidimensional Computerized Adaptive Testing
British Journal of Mathematical and Statistical Psychology, 2017Co-Authors: Ping Chen, Tao Xin, Chun Wang, Huahua ChangAbstract:Multidimensional Computerized Adaptive Testing (MCAT) has received increasing attention over the past few years in educational measurement. Like all other formats of CAT, item replenishment is an essential part of MCAT for its item bank maintenance and management, which governs retiring overexposed or obsolete items over time and replacing them with new ones. Moreover, calibration precision of the new items will directly affect the estimation accuracy of examinees' ability vectors. In unidimensional CAT (UCAT) and cognitive diagnostic CAT, online calibration techniques have been developed to effectively calibrate new items. However, there has been very little discussion of online calibration in MCAT in the literature. Thus, this paper proposes new online calibration methods for MCAT based upon some popular methods used in UCAT. Three representative methods, Method A, the 'one EM cycle' method and the 'multiple EM cycles' method, are generalized to MCAT. Three simulation studies were conducted to compare the three new methods by manipulating three factors (test length, item bank design, and level of correlation between coordinate dimensions). The results showed that all the new methods were able to recover the item parameters accurately, and the Adaptive online calibration designs showed some improvements compared to the random design under most conditions.
Juan Ramon Barrada - One of the best experts on this subject based on the ideXlab platform.
-
adapting cognitive diagnosis Computerized Adaptive Testing item selection rules to traditional item response theory
PLOS ONE, 2020Co-Authors: Miguel A. Sorrel, Jimmy De La Torre, Juan Ramon Barrada, Francisco J AbadAbstract:: Currently, there are two predominant approaches in Adaptive Testing. One, referred to as cognitive diagnosis Computerized Adaptive Testing (CD-CAT), is based on cognitive diagnosis models, and the other, the traditional CAT, is based on item response theory. The present study evaluates the performance of two item selection rules (ISRs) originally developed in the CD-CAT framework, the double Kullback-Leibler information (DKL) and the generalized deterministic inputs, noisy "and" gate model discrimination index (GDI), in the context of traditional CAT. The accuracy and test security associated with these two ISRs are compared to those of the point Fisher information and weighted KL using a simulation study. The impact of the trait level estimation method is also investigated. The results show that the new ISRs, particularly DKL, could be used to improve the accuracy of CAT. Better accuracy for DKL is achieved at the expense of higher item overlap rate. Differences among the item selection rules become smaller as the test gets longer. The two CD-CAT ISRs select different types of items: items with the highest possible a parameter with DKL, and items with the lowest possible c parameter with GDI. Regarding the trait level estimator, expected a posteriori method is generally better in the first stages of the CAT, and converges with the maximum likelihood method when a medium to large number of items are involved. The use of DKL can be recommended in low-stakes settings where test security is less of a concern.
-
Computerized Adaptive Testing with r recent updates of the package catr
Journal of Statistical Software, 2017Co-Authors: David Magis, Juan Ramon BarradaAbstract:The purpose of this paper is to list the recent updates of the R package catR. This package allows for generating response patterns under a Computerized Adaptive Testing (CAT) framework with underlying item response theory (IRT) models. Among the most important updates, well-known polytomous IRT models are now supported by catR; several item selection rules have been added; and it is now possible to perform post-hoc simulations. Some functions were also rewritten or withdrawn to improve the usefulness and performances of the package.
-
new item selection methods for cognitive diagnosis Computerized Adaptive Testing
Applied Psychological Measurement, 2015Co-Authors: Mehmet Kaplan, Jimmy De La Torre, Juan Ramon BarradaAbstract:This article introduces two new item selection methods, the modified posterior-weighted Kullback-Leibler index (MPWKL) and the generalized deterministic inputs, noisy "and" gate (G-DINA) model discrimination index (GDI), that can be used in cognitive diagnosis Computerized Adaptive Testing. The efficiency of the new methods is compared with the posterior-weighted Kullback-Leibler (PWKL) item selection index using a simulation study in the context of the G-DINA model. The impact of item quality, generating models, and test termination rules on attribute classification accuracy or test length is also investigated. The results of the study show that the MPWKL and GDI perform very similarly, and have higher correct attribute classification rates or shorter mean test lengths compared with the PWKL. In addition, the GDI has the shortest implementation time among the three indices. The proportion of item usage with respect to the required attributes across the different conditions is also tracked and discussed.
Zhuoran Shang - One of the best experts on this subject based on the ideXlab platform.
-
variable length stopping rules for multidimensional Computerized Adaptive Testing
Psychometrika, 2019Co-Authors: Chun Wang, David J Weiss, Zhuoran ShangAbstract:In Computerized Adaptive Testing (CAT), a variable-length stopping rule refers to ending item administration after a pre-specified measurement precision standard has been satisfied. The goal is to provide equal measurement precision for all examinees regardless of their true latent trait level. Several stopping rules have been proposed in unidimensional CAT, such as the minimum information rule or the maximum standard error rule. These rules have also been extended to multidimensional CAT and cognitive diagnostic CAT, and they all share the same idea of monitoring measurement error. Recently, Babcock and Weiss (J Comput Adapt Test 2012. https://doi.org/10.7333/1212-0101001) proposed an “absolute change in theta” (CT) rule, which is useful when an item bank is exhaustive of good items for one or more ranges of the trait continuum. Choi, Grady and Dodd (Educ Psychol Meas 70:1–17, 2010) also argued that a CAT should stop when the standard error does not change, implying that the item bank is likely exhausted. Although these stopping rules have been evaluated and compared in different simulation studies, the relationships among the various rules remain unclear, and therefore there lacks a clear guideline regarding when to use which rule. This paper presents analytic results to show the connections among various stopping rules within both unidimensional and multidimensional CAT. In particular, it is argued that the CT-rule alone can be unstable and it can end the test prematurely. However, the CT-rule can be a useful secondary rule to monitor the point of diminished returns. To further provide empirical evidence, three simulation studies are reported using both the 2PL model and the multidimensional graded response model.
-
on initial item selection in cognitive diagnostic Computerized Adaptive Testing
British Journal of Mathematical and Statistical Psychology, 2016Co-Authors: Chun Wang, Zhuoran ShangAbstract:There has recently been much interest in Computerized Adaptive Testing (CAT) for cognitive diagnosis. While there exist various item selection criteria and different asymptotically optimal designs, these are mostly constructed based on the asymptotic theory assuming the test length goes to infinity. In practice, with limited test lengths, the desired asymptotic optimality may not always apply, and there are few studies in the literature concerning the optimal design of finite items. Related questions, such as how many items we need in order to be able to identify the attribute pattern of an examinee and what types of initial items provide the optimal classification results, are still open. This paper aims to answer these questions by providing non-asymptotic theory of the optimal selection of initial items in cognitive diagnostic CAT. In particular, for the optimal design, we provide necessary and sufficient conditions for the Q-matrix structure of the initial items. The theoretical development is suitable for a general family of cognitive diagnostic models. The results not only provide a guideline for the design of optimal item selection procedures, but also may be applied to guide item bank construction.