Skip Navigation
Skip to contents

JEEHP : Journal of Educational Evaluation for Health Professions

OPEN ACCESS
SEARCH
Search

Search

Page Path
HOME > Search
7 "Jeongwook Choi"
Filter
Filter
Article category
Keywords
Publication year
Authors
Funded articles
Research articles
Psychometric characteristics of Korean Medical Licensing Examination items from 2012 to 2022 analyzed using the Rasch model and 2-parameter logistic model
Dong Gi Seo, Jeongwook Choi, Mee Young Kim, Sun Huh
J Educ Eval Health Prof. 2026;23:23.   Published online August 4, 2026
DOI: https://doi.org/10.3352/jeehp.2026.23.23    [Epub ahead of print]
  • 391 View
  • 49 Download
AbstractAbstract PDF
Purpose
This study characterized the psychometric properties of Korean Medical Licensing Examination items administered from 2012 to 2022 using the Rasch and 2-parameter logistic models and descriptively examined items classified as highly difficult.
Methods
Item parameters were estimated separately for each examination year using the Rasch and 2-parameter logistic models implemented in the irtQ package in R. Descriptive statistics and correlations between model-based difficulty estimates were calculated. Items with a 2-parameter logistic difficulty estimate of b ≥2.0 were subjected to content analysis.
Results
Correlations between Rasch and 2-parameter logistic item-difficulty estimates ranged from 0.70 to 0.80 across examination years. Under the 2-parameter logistic model, the proportion of items with b ≥2.0 ranged from 4.7% to 8.9% and showed no monotonic temporal trend; the proportions were 5.0% in 2021 and 5.6% in 2022. The Rasch model generally produced smaller estimated standard errors than the 2-parameter logistic model.
Conclusion
Rasch and 2-parameter logistic difficulty estimates showed strongly correlated rankings, although correlation alone did not establish agreement between their absolute estimates. The Rasch model demonstrated greater numerical stability under the present calibration conditions, but model selection should also consider model fit, test information, and the intended use of scores. Because examination years were calibrated separately and no pre-disclosure or control condition was available, the observed annual differences cannot be interpreted as evidence that item disclosure caused changes in item difficulty or discrimination.
The impact of differential item functioning on ability estimation using the Korean Medical Licensing Examination with computerized adaptive testing: a post-hoc simulation study  
Dogyeong Kim, Jeongwook Choi, Dong Gi Seo
J Educ Eval Health Prof. 2025;22:31.   Published online October 10, 2025
DOI: https://doi.org/10.3352/jeehp.2025.22.31
  • 3,805 View
  • 200 Download
AbstractAbstract PDFSupplementary Material
Purpose
This study examined the impact of differential item functioning (DIF) on ability estimation in a computerized adaptive testing (CAT) environment using real response data from the 2017 Korean Medical Licensing Examination (KMLE). We hypothesized that excluding gender-based DIF items would improve estimation accuracy, particularly for examinees at the extremes of the ability scale.
Methods
The study was conducted in 2 steps: (1) DIF detection and (2) post-hoc simulation. The analysis used data from 3,259 examinees who completed all 360 dichotomous items. Gender-based DIF was detected with the residual-based DIF method (reference group: males; focal group: females). Two CAT conditions (all items vs. DIF-excluded) were compared against a “true θ” estimated from a fixed-form test of 264 non-DIF items. Accuracy was evaluated using bias, root mean square error (RMSE), and correlation with true θ.
Results
In the CAT condition excluding DIF items, accuracy improved, with RMSE reduced and correlation with true θ increased. However, bias was slightly larger in magnitude. Gender-specific analyses showed that DIF removal reduced the underestimation of female ability but increased the underestimation of male ability, yielding estimates that were fairer across genders. When DIF items were included, estimation errors were more pronounced at both low and high ability levels.
Conclusion
Managing DIF in CAT-based high-stakes examinations can enhance fairness and precision. Using real examinee data, this study provides practical evidence of the implications of DIF for CAT-based measurement and supports fairness-oriented test design.
Technical report
Feasibility of applying computerized adaptive testing to the Clinical Medical Science Comprehensive Examination in Korea: a psychometric study  
Jeongwook Choi, Sung-Soo Jung, Eun Kwang Choi, Kyung Sik Kim, Dong Gi Seo
J Educ Eval Health Prof. 2025;22:29.   Published online October 1, 2025
DOI: https://doi.org/10.3352/jeehp.2025.22.29
  • 1,980 View
  • 216 Download
AbstractAbstract PDFSupplementary Material
Purpose
This study aimed to investigate the feasibility of transitioning the Clinical Medical Science Comprehensive Examination (CMSCE) to computerized adaptive testing (CAT) in Korea, thereby providing greater opportunities for medical students to accurately compare their clinical competencies with peers nationwide and to monitor their own progress.
Methods
A medical self-assessment using CAT was conducted from March to June 2023, involving 1,541 medical students who volunteered from 40 medical colleges in Korea. An item bank consisting of 1,145 items from previously administered CMSCE examinations (2019–2021) hosted by the Medical Education Assessment Corporation was established. Items were selected through 2-stage filtering, based on classical test theory (discrimination index above 0.15) and item response theory (discrimination parameter estimates above 0.6 and difficulty parameter estimates between –5 and +5). Maximum Fisher information was employed as the item selection method, and maximum likelihood estimation was used for ability estimation.
Results
The CAT was successfully administered without significant issues. The stopping rule was set at a standard error of measurement of 0.25, with a maximum of 50 items for ability estimation. The mean ability score was 0.55, with an average of 28 items administered per student. Students at extreme ability levels reached the maximum of 50 items due to the limited availability of items at appropriate difficulty levels.
Conclusion
The medical self-assessment CAT, the first of its kind in Korea, was successfully implemented nationwide without significant problems. These results indicate strong potential for expanding the use of CAT in medical education assessments.
Research article
Special article on the 20th anniversary of the journal
Comparison of real data and simulated data analysis of a stopping rule based on the standard error of measurement in computerized adaptive testing for medical examinations in Korea: a psychometric study  
Dong Gi Seo, Jeongwook Choi, Jinha Kim
J Educ Eval Health Prof. 2024;21:18.   Published online July 9, 2024
DOI: https://doi.org/10.3352/jeehp.2024.21.18
  • 4,750 View
  • 386 Download
  • 2 Web of Science
  • 2 Crossref
AbstractAbstract PDFSupplementary Material
Purpose
This study aimed to compare and evaluate the efficiency and accuracy of computerized adaptive testing (CAT) under 2 stopping rules (standard error of measurement [SEM]=0.3 and 0.25) using both real and simulated data in medical examinations in Korea.
Methods
This study employed post-hoc simulation and real data analysis to explore the optimal stopping rule for CAT in medical examinations. The real data were obtained from the responses of 3rd-year medical students during examinations in 2020 at Hallym University College of Medicine. Simulated data were generated using estimated parameters from a real item bank in R. Outcome variables included the number of examinees’ passing or failing with SEM values of 0.25 and 0.30, the number of items administered, and the correlation. The consistency of real CAT result was evaluated by examining consistency of pass or fail based on a cut score of 0.0. The efficiency of all CAT designs was assessed by comparing the average number of items administered under both stopping rules.
Results
Both SEM 0.25 and SEM 0.30 provided a good balance between accuracy and efficiency in CAT. The real data showed minimal differences in pass/fail outcomes between the 2 SEM conditions, with a high correlation (r=0.99) between ability estimates. The simulation results confirmed these findings, indicating similar average item numbers between real and simulated data.
Conclusion
The findings suggest that both SEM 0.25 and 0.30 are effective termination criteria in the context of the Rasch model, balancing accuracy and efficiency in CAT.

Citations

Citations to this article as recorded by  
  • AI-enhanced adaptive testing with cognitive diagnostic feedback and its association with performance in undergraduate surgical education: a pilot study
    Nuno Silva Gonçalves, Carlos Collares, José Miguel Pêgo
    Frontiers in Behavioral Neuroscience.2026;[Epub]     CrossRef
  • Feasibility of applying computerized adaptive testing to the Clinical Medical Science Comprehensive Examination in Korea: a psychometric study
    Jeongwook Choi, Sung-Soo Jung, Eun Kwang Choi, Kyung Sik Kim, Dong Gi Seo
    Journal of Educational Evaluation for Health Professions.2025; 22: 29.     CrossRef
Software report
Introduction to the LIVECAT web-based computerized adaptive testing platform  
Dong Gi Seo, Jeongwook Choi
J Educ Eval Health Prof. 2020;17:27.   Published online September 29, 2020
DOI: https://doi.org/10.3352/jeehp.2020.17.27
  • 9,373 View
  • 164 Download
  • 7 Web of Science
  • 10 Crossref
AbstractAbstract PDFSupplementary Material
This study introduces LIVECAT, a web-based computerized adaptive testing platform. This platform provides many functions, including writing item content, managing an item bank, creating and administering a test, reporting test results, and providing information about a test and examinees. The LIVECAT provides examination administrators with an easy and flexible environment for composing and managing examinations. It is available at http://www.thecatkorea.com/. Several tools were used to program LIVECAT, as follows: operating system, Amazon Linux; web server, nginx 1.18; WAS, Apache Tomcat 8.5; database, Amazon RDMS—Maria DB; and languages, JAVA8, HTML5/CSS, Javascript, and jQuery. The LIVECAT platform can be used to implement several item response theory (IRT) models such as the Rasch and 1-, 2-, 3-parameter logistic models. The administrator can choose a specific model of test construction in LIVECAT. Multimedia data such as images, audio files, and movies can be uploaded to items in LIVECAT. Two scoring methods (maximum likelihood estimation and expected a posteriori) are available in LIVECAT and the maximum Fisher information item selection method is applied to every IRT model in LIVECAT. The LIVECAT platform showed equal or better performance compared with a conventional test platform. The LIVECAT platform enables users without psychometric expertise to easily implement and perform computerized adaptive testing at their institutions. The most recent LIVECAT version only provides a dichotomous item response model and the basic components of CAT. Shortly, LIVECAT will include advanced functions, such as polytomous item response models, weighted likelihood estimation method, and content balancing method.

Citations

Citations to this article as recorded by  
  • Development and Evaluation of Computerized Adaptive Testing for General English Proficiency Test
    Hyeon Park, Woojae Han, Mikyung Park, Dong gi Seo
    The Korean Association of General Education.2026; 20(2): 191.     CrossRef
  • A Systematic Review on Computerized Adaptive Testing
    Hümeyra Demir, Selahattin Gelbal
    Erzincan Üniversitesi Eğitim Fakültesi Dergisi.2025; 27(1): 137.     CrossRef
  • Development of a CAT based Diagnostic System for Assessing Basic Academic Skills in Undergraduate Students
    Woo-Jin Han, Jeongwook Choi, Dong-Gi Seo
    The Korean Association of General Education.2025; 19(3): 177.     CrossRef
  • Feasibility of applying computerized adaptive testing to the Clinical Medical Science Comprehensive Examination in Korea: a psychometric study
    Jeongwook Choi, Sung-Soo Jung, Eun Kwang Choi, Kyung Sik Kim, Dong Gi Seo
    Journal of Educational Evaluation for Health Professions.2025; 22: 29.     CrossRef
  • Comparison of real data and simulated data analysis of a stopping rule based on the standard error of measurement in computerized adaptive testing for medical examinations in Korea: a psychometric study
    Dong Gi Seo, Jeongwook Choi, Jinha Kim
    Journal of Educational Evaluation for Health Professions.2024; 21: 18.     CrossRef
  • Educational Technology in the University: A Comprehensive Look at the Role of a Professor and Artificial Intelligence
    Cheolkyu Shin, Dong Gi Seo, Seoyeon Jin, Soo Hwa Lee, Hyun Je Park
    IEEE Access.2024; 12: 116727.     CrossRef
  • The irtQ R package: a user-friendly tool for item response theory-based test data analysis and calibration
    Hwanggyu Lim, Kyungseok Kang
    Journal of Educational Evaluation for Health Professions.2024; 21: 23.     CrossRef
  • Presidential address: improving item validity and adopting computer-based testing, clinical skills assessments, artificial intelligence, and virtual reality in health professions licensing examinations in Korea
    Hyunjoo Pai
    Journal of Educational Evaluation for Health Professions.2023; 20: 8.     CrossRef
  • Patient-reported outcome measures in cancer care: Integration with computerized adaptive testing
    Minyu Liang, Zengjie Ye
    Asia-Pacific Journal of Oncology Nursing.2023; 10(12): 100323.     CrossRef
  • Development of a character qualities test for medical students in Korea using polytomous item response theory and factor analysis: a preliminary scale development study
    Yera Hur, Dong Gi Seo
    Journal of Educational Evaluation for Health Professions.2023; 20: 20.     CrossRef
Corrigendum
Funding information of the article entitled “Post-hoc simulation study of computerized adaptive testing for the Korean Medical Licensing Examination”
Dong Gi Seo, Jeongwook Choi
J Educ Eval Health Prof. 2018;15:27.   Published online December 4, 2018
DOI: https://doi.org/10.3352/jeehp.2018.15.27    [Epub ahead of print]
Corrects: J Educ Eval Health Prof 2018;15(0):14
  • 19,339 View
  • 192 Download
PDF
Research article
Post-hoc simulation study of computerized adaptive testing for the Korean Medical Licensing Examination  
Dong Gi Seo, Jeongwook Choi
J Educ Eval Health Prof. 2018;15:14.   Published online May 17, 2018
DOI: https://doi.org/10.3352/jeehp.2018.15.14
Correction in: J Educ Eval Health Prof 2018;15(0):27
  • 40,460 View
  • 362 Download
  • 11 Web of Science
  • 13 Crossref
AbstractAbstract PDFSupplementary Material
Purpose
Computerized adaptive testing (CAT) has been adopted in licensing examinations because it improves the efficiency and accuracy of the tests, as shown in many studies. This simulation study investigated CAT scoring and item selection methods for the Korean Medical Licensing Examination (KMLE).
Methods
This study used a post-hoc (real data) simulation design. The item bank used in this study included all items from the January 2017 KMLE. All CAT algorithms for this study were implemented using the ‘catR’ package in the R program.
Results
In terms of accuracy, the Rasch and 2-parametric logistic (PL) models performed better than the 3PL model. The ‘modal a posteriori’ and ‘expected a posterior’ methods provided more accurate estimates than maximum likelihood estimation or weighted likelihood estimation. Furthermore, maximum posterior weighted information and minimum expected posterior variance performed better than other item selection methods. In terms of efficiency, the Rasch model is recommended to reduce test length.
Conclusion
Before implementing live CAT, a simulation study should be performed under varied test conditions. Based on a simulation study, and based on the results, specific scoring and item selection methods should be predetermined.

Citations

Citations to this article as recorded by  
  • Development of a Brief Computerized Adaptive Testing Version of the SCL-90: Validation of a Six-Factor Model Using a Chinese Adolescent Clinical Sample
    Yuanlu Xiong, Hao Xu, Jiaqi Song, Yan Chen, Zhiren Wang, Ronghuan Jiang, Shuping Tan
    Journal of Personality Assessment.2026; : 1.     CrossRef
  • Large-Scale Parallel Cognitive Diagnostic Test Assembly Using A Dual-Stage Differential Evolution-Based Approach
    Xi Cao, Ying Lin, Dong Liu, Henry Been-Lirn Duh, Jun Zhang
    IEEE Transactions on Artificial Intelligence.2024; 5(6): 3120.     CrossRef
  • Assessing the Potentials of Compurized Adaptive Testing to Enhance Mathematics and Science Student’t Achievement in Secondary Schools
    Mary Patrick Uko, I.O. Eluwa, Patrick J. Uko
    European Journal of Theoretical and Applied Sciences.2024; 2(4): 85.     CrossRef
  • Comparison of real data and simulated data analysis of a stopping rule based on the standard error of measurement in computerized adaptive testing for medical examinations in Korea: a psychometric study
    Dong Gi Seo, Jeongwook Choi, Jinha Kim
    Journal of Educational Evaluation for Health Professions.2024; 21: 18.     CrossRef
  • Comparison of CAT Procedures at Low Ability Levels: A Simulation Study and Analysis in the Context of Students with Disabilities
    Selma Şenel
    Bartın University Journal of Faculty of Education.2024; 13(3): 547.     CrossRef
  • Presidential address: improving item validity and adopting computer-based testing, clinical skills assessments, artificial intelligence, and virtual reality in health professions licensing examinations in Korea
    Hyunjoo Pai
    Journal of Educational Evaluation for Health Professions.2023; 20: 8.     CrossRef
  • Developing Computerized Adaptive Testing for a National Health Professionals Exam: An Attempt from Psychometric Simulations
    Lingling Xu, Zhehan Jiang, Yuting Han, Haiying Liang, Jinying Ouyang
    Perspectives on Medical Education.2023;[Epub]     CrossRef
  • Optimizing Computer Adaptive Test Performance: A Hybrid Simulation Study to Customize the Administration Rules of the CAT-EyeQ in Macular Edema Patients
    T. Petra Rausch-Koster, Michiel A. J. Luijten, Frank D. Verbraak, Ger H. M. B. van Rens, Ruth M. A. van Nispen
    Translational Vision Science & Technology.2022; 11(11): 14.     CrossRef
  • The accuracy and consistency of mastery for each content domain using the Rasch and deterministic inputs, noisy “and” gate diagnostic classification models: a simulation study and a real-world analysis using data from the Korean Medical Licensing Examinat
    Dong Gi Seo, Jae Kum Kim
    Journal of Educational Evaluation for Health Professions.2021; 18: 15.     CrossRef
  • A Seed Usage Issue on Using catR for Simulation and the Solution
    Zhongmin Cui
    Applied Psychological Measurement.2020; 44(5): 409.     CrossRef
  • Linear programming method to construct equated item sets for the implementation of periodical computer-based testing for the Korean Medical Licensing Examination
    Dong Gi Seo, Myeong Gi Kim, Na Hui Kim, Hye Sook Shin, Hyun Jung Kim
    Journal of Educational Evaluation for Health Professions.2018; 15: 26.     CrossRef
  • Funding information of the article entitled “Post-hoc simulation study of computerized adaptive testing for the Korean Medical Licensing Examination”
    Dong Gi Seo, Jeongwook Choi
    Journal of Educational Evaluation for Health Professions.2018; 15: 27.     CrossRef
  • Updates from 2018: Being indexed in Embase, becoming an affiliated journal of the World Federation for Medical Education, implementing an optional open data policy, adopting principles of transparency and best practice in scholarly publishing, and appreci
    Sun Huh
    Journal of Educational Evaluation for Health Professions.2018; 15: 36.     CrossRef

JEEHP : Journal of Educational Evaluation for Health Professions
TOP