Skip Navigation
Skip to contents

JEEHP : Journal of Educational Evaluation for Health Professions

OPEN ACCESS
SEARCH
Search

Search

Page Path
HOME > Search
216 "Education"
Filter
Filter
Article category
Keywords
Publication year
Authors
Funded articles
Research articles
Refusal-aware evaluation of frontier AI models available in June 2026 using the Japanese National License Examination for Pharmacists: a comparative study
Hiroyasu Sato, Katsuhiko Ogasawara, Hidehiko Sakurai
J Educ Eval Health Prof. 2026;23:29.   Published online September 3, 2026
DOI: https://doi.org/10.3352/jeehp.2026.23.29    [Epub ahead of print]
  • 223 View
  • 23 Download
AbstractAbstract PDF
Purpose
Conventional single-run accuracy may be insufficient for evaluating frontier generative artificial intelligence (AI) models when safety-related refusals occur. This 4-model benchmark examined the need for repeated, refusal-aware evaluation using the Japanese National License Examination for Pharmacists (JNLEP).
Methods
ChatGPT GPT-5.5, Gemini 3.5 Flash, Claude Opus 4.8, and Claude Fable 5 were evaluated using all 345 questions from the 107th JNLEP. The original Japanese questions, including image-containing items, were submitted through application programming interfaces (APIs) in 3 independent runs. Refusals were treated as incorrect when overall accuracy was calculated. For Fable 5, accuracy excluding refusals, refusal rate, refusal consistency across runs, the subject-wise distribution of refusals, and system-assigned refusal categories were also evaluated.
Results
The mean overall accuracies were 98.7% for GPT-5.5, 98.3% for Gemini 3.5 Flash, 96.1% for Claude Opus 4.8, and 70.3% for Claude Fable 5. Fable 5 had a mean refusal rate of 29.0%, whereas its mean accuracy excluding refusals was 99.0%. Among the 345 items, 93 were refused in all 3 runs, 14 were refused inconsistently across runs, and 238 were never refused. All 300 refusal responses were assigned to the bio category. Refusals were most frequent in Biology (90.0%) and Pharmacology (69.2%) but uncommon in Practice (2.5%).
Conclusion
Near-saturation benchmark performance coexisted with frequent and partly run-dependent refusals in a safeguard-equipped model. Overall accuracy, accuracy excluding refusals, refusal rate, and refusal consistency describe complementary aspects of performance; therefore, repeated, refusal-aware evaluation is needed to interpret frontier AI models in pharmacy education.
Development and initial validation of an instrument measuring the perceived importance of core competencies for clinical nurse educators in Korea: a methodological study  
Heui-Kyeong Kwon, Junmoo Ahn, Kyungsan Seo, Seung-Hee Lee, Eunhee Jung, Yoonjung Lee
J Educ Eval Health Prof. 2026;23:28.   Published online September 3, 2026
DOI: https://doi.org/10.3352/jeehp.2026.23.28    [Epub ahead of print]
  • 191 View
  • 23 Download
AbstractAbstract PDFSupplementary Material
Purpose
This study aimed to develop an instrument to assess the perceived importance of core competencies for clinical nurse educators (CNEs) in Korea, defined here as hospital-based nurse educators under the Korean Nursing Act, and to provide initial evidence of its content validity, internal structure, and reliability.
Methods
A preliminary 44-item pool, developed from international nurse educator frameworks and the literature, was refined by 8 experts through a 2-round modified Delphi process. A nationwide sample of 263 CNEs then rated the perceived importance of each item. Dimensionality was examined using parallel analysis and exploratory factor analysis based on polychoric correlations; the refined model was tested using confirmatory factor analysis (CFA) with the weighted least squares mean- and variance-adjusted estimator, composite reliability (CR), average variance extracted (AVE), Fornell-Larcker and heterotrait-monotrait discriminant validity, and alternative models.
Results
Content validity was high (scale-level content validity index=0.986). Parallel analysis supported a 5-factor structure; 9 weak or redundant items were removed, consolidating the 8 domains into 5 factors comprising 35 items. The CFA model showed good fit (comparative fit index=0.977, Tucker-Lewis index=0.975, root mean square error of approximation=0.048, standardized root mean square residual=0.053; standardized loadings, 0.669–0.934). Internal consistency was high (Cronbach’s α=0.873–0.913; categorical McDonald’s ω=0.871–0.916), and convergent validity was supported (CR=0.91–0.95; AVE>0.50). Three factors were highly correlated; a second-order model fit well, and the general factor explained most of the reliable variance (ω_h=0.83).
Conclusion
The 35-item instrument provided initial evidence of content validity, internal structure, and reliability for assessing the perceived importance of CNE competencies. These findings represent initial, rather than definitive, validity evidence and require cross-validation in an independent sample.
Name-based instructor–student interaction in medical teaching: development and psychometric evaluation of a pilot questionnaire  
Liam Philipp Weber, Philipp Huber, Julius Wiemschulte
J Educ Eval Health Prof. 2026;23:26.   Published online August 13, 2026
DOI: https://doi.org/10.3352/jeehp.2026.23.26    [Epub ahead of print]
  • 550 View
  • 85 Download
AbstractAbstract PDF
Purpose
When instructors address students by name, this may be associated with students’ engagement and sense of inclusion in higher education. However, name-based instructor–student interaction has rarely been conceptualized as a multidimensional construct, and no psychometrically evaluated instrument is available for medical education. This study developed and psychometrically evaluated a pilot questionnaire assessing students’ perceived experiences of being addressed by name in medical teaching contexts.
Methods
Questionnaire development followed a multistep exploratory approach that included qualitative input, iterative item refinement, and expert content review. An initial pool of 59 items was administered to medical students (n=270) from semesters 2–10 at a German medical faculty. Factor structure was examined using exploratory factor analysis based on Pearson correlations, oblique rotation, and predefined item-reduction criteria. Internal consistency was assessed using McDonald’s ω and Cronbach’s α, discriminant validity using the heterotrait–monotrait ratio, and generalized partial credit models were estimated separately for each factor.
Results
Exploratory factor analysis yielded a parsimonious 3-factor solution comprising 15 items: cognitive activation, appreciation & social belonging, and social evaluative anxiety. The final model showed good fit, and all factors demonstrated strong internal consistency (ω=0.84–0.88), good measurement precision (expected a posteriori reliability=0.83–0.88), adequate discriminant validity (heterotrait–monotrait ratio <0.85), and acceptable item fit.
Conclusion
The pilot questionnaire captures 3 distinct dimensions of name-based instructor–student interaction and provides promising initial psychometric evidence for medical education research. It may support confirmatory validation and further study of relational teaching practices.
Short-term knowledge gains and learner satisfaction after a basic radiology course for family medicine residents in a 3-dimensional virtual environment: a multi-cohort pretest–posttest study in Spain  
Alberto Pino-Postigo, Rocio Lorenzo-Alvarez, Dolores Dominguez-Pinos, Eugenio Lázaro Navarro-Sanchis, Eduardo Ochando-Pulido, Miguel Jose-Ruiz Gomez, Francisco Sendra-Portero
J Educ Eval Health Prof. 2026;23:25.   Published online August 7, 2026
DOI: https://doi.org/10.3352/jeehp.2026.23.25    [Epub ahead of print]
  • 458 View
  • 68 Download
AbstractAbstract PDF
Purpose
This study aimed to evaluate a structured basic radiology course for Family Medicine residents delivered in a 3-dimensional (3D) virtual environment (Second Life). We hypothesized that participation would be associated with higher post-course knowledge test scores and high learner satisfaction.
Methods
A quasi-experimental multi-cohort pretest–posttest study was conducted across 3 consecutive cohorts of Family Medicine residents in Spain in 2019. Ninety-six participants engaged in a 15-day course combining synchronous and asynchronous activities. Sixty-five participants provided paired pre- and post-intervention knowledge assessments, which were analyzed using paired t-tests. Learner satisfaction was evaluated using a structured questionnaire with Likert-scale items, numerical ratings (0–10), and open-ended responses. Pre- and post-test scores and ratings were compared using paired t-tests, and Mann-Whitney U tests were used for Likert-scale data. Qualitative data from open-ended responses were analyzed using thematic coding.
Results
Participants had significantly higher post-course knowledge test scores, increasing from 47.4±11.6 to 58.2±11.9 (mean difference, 10.8; P<0.001; paired-sample Cohen’s d [dz]=0.67). High agreement was observed across most satisfaction items, with median scores ranging from 4 to 5. Content relevance, usefulness for clinical practice, and instructor performance received the highest ratings. Lower scores were observed for peer interaction and platform usability.
Conclusion
A structured radiology course delivered in a 3D virtual environment was associated with higher immediate post-course knowledge test scores and high learner satisfaction among Family Medicine residents. The findings support the feasibility of repeated course delivery, although controlled studies are warranted.
Review
The impact of the learner role in healthcare simulation-based education on learning outcomes: a systematic review
Abdelkader Amechghal, Aziz Naciri, Kamal Takhdat, Sana Loubbairi, Adil Mellouki, Laila Lahlou, Hicham Nassik
J Educ Eval Health Prof. 2026;23:24.   Published online August 6, 2026
DOI: https://doi.org/10.3352/jeehp.2026.23.24    [Epub ahead of print]
  • 594 View
  • 93 Download
AbstractAbstract PDF
Purpose
This review examined the impact of learner role—active participant versus observer—on learning outcomes in healthcare simulation-based education using Kirkpatrick’s evaluation model.
Methods
Five databases—PubMed, Scopus, Web of Science, ScienceDirect, and EBSCOhost—were searched for studies published from January 2018 to November 2024. Eligibility criteria, defined using the PICOS framework, targeted studies comparing learning outcomes between active participant and observer roles among health professions students and practitioners. Methodological quality was appraised using the Medical Education Research Study Quality Instrument, and findings were synthesized narratively. This review was registered in PROSPERO under registration number CRD42024611988.
Results
Of 648 records, 15 studies involving 1,799 participants met the inclusion criteria; role allocation was specifically reported for 644 active participants and 709 observers. Most studies were conducted in high-income countries (13/15) and nursing programs (11/15). At Kirkpatrick Level 1, most studies reported no statistically significant between-role differences, although an active participant advantage emerged for emotional arousal. At Level 2, knowledge, skills, and most attitudinal outcomes did not differ significantly between roles; active participant advantages were observed for technical skills at delayed follow-up, affective competencies in palliative care, and perceived learning transfer. No study assessed Levels 3 or 4.
Conclusion
Most studies found no significant differences between observer and active participant roles at Kirkpatrick Levels 1 and 2. Generalizability across all health professions and settings is limited. Pedagogically appropriate use of the observer role supports learning outcomes and may expand access to simulation-based education.
Research articles
Psychometric characteristics of Korean Medical Licensing Examination items from 2012 to 2022 analyzed using the Rasch model and 2-parameter logistic model
Dong Gi Seo, Jeongwook Choi, Mee Young Kim, Sun Huh
J Educ Eval Health Prof. 2026;23:23.   Published online August 4, 2026
DOI: https://doi.org/10.3352/jeehp.2026.23.23    [Epub ahead of print]
  • 391 View
  • 49 Download
AbstractAbstract PDF
Purpose
This study characterized the psychometric properties of Korean Medical Licensing Examination items administered from 2012 to 2022 using the Rasch and 2-parameter logistic models and descriptively examined items classified as highly difficult.
Methods
Item parameters were estimated separately for each examination year using the Rasch and 2-parameter logistic models implemented in the irtQ package in R. Descriptive statistics and correlations between model-based difficulty estimates were calculated. Items with a 2-parameter logistic difficulty estimate of b ≥2.0 were subjected to content analysis.
Results
Correlations between Rasch and 2-parameter logistic item-difficulty estimates ranged from 0.70 to 0.80 across examination years. Under the 2-parameter logistic model, the proportion of items with b ≥2.0 ranged from 4.7% to 8.9% and showed no monotonic temporal trend; the proportions were 5.0% in 2021 and 5.6% in 2022. The Rasch model generally produced smaller estimated standard errors than the 2-parameter logistic model.
Conclusion
Rasch and 2-parameter logistic difficulty estimates showed strongly correlated rankings, although correlation alone did not establish agreement between their absolute estimates. The Rasch model demonstrated greater numerical stability under the present calibration conditions, but model selection should also consider model fit, test information, and the intended use of scores. Because examination years were calibrated separately and no pre-disclosure or control condition was available, the observed annual differences cannot be interpreted as evidence that item disclosure caused changes in item difficulty or discrimination.
Development and psychometric validation of the Thai AI Literacy Scale for Nursing Students in Thailand: a methodological study
Wannaporn Jongchidklang, Supalak Phonphithak, Porntep Amornritvanich
J Educ Eval Health Prof. 2026;23:22.   Published online August 4, 2026
DOI: https://doi.org/10.3352/jeehp.2026.23.22    [Epub ahead of print]
  • 707 View
  • 109 Download
AbstractAbstract PDF
Purpose
This study aimed to develop and validate the Thai AI Literacy Scale for Nursing Students (TAILS-NS).
Methods
A cross-sectional study was conducted across multiple nursing institutions in Thailand from March to May 2026 to address the absence of a validated artificial intelligence (AI) literacy instrument for nursing students in Thai or Southeast Asian contexts. A total of 410 nursing students participated, yielding a response rate of 94.5%. The TAILS-NS was developed through item generation, expert content-validity assessment using the item-objective congruence index, and a 2-phase pilot study, resulting in a 40-item instrument comprising a 15-item knowledge test and 25 Likert-scale items across 6 domains. Exploratory factor analysis using maximum likelihood estimation and confirmatory factor analysis using the weighted least squares mean and variance-adjusted estimator, as well as internal consistency, convergent validity, and discriminant validity, were assessed. An independent validation sample (n=157) was recruited for cross-validation confirmatory factor analysis. Raw data are available as a supplement.
Results
Exploratory factor analysis supported a 6-factor structure: AI awareness, skills, ethics and professionalism, positive attitude, AI anxiety, and readiness. Confirmatory factor analysis showed acceptable model fit, although the root mean square error of approximation (RMSEA) indicated marginal fit (comparative fit index [CFI]=0.919, Tucker-Lewis index [TLI]=0.906, RMSEA=0.090). Reliability of the knowledge subscale was acceptable (Kuder-Richardson Formula 20=0.758). Internal consistency was excellent (α=0.810–0.934; total α=0.974; ω=0.989). Average variance extracted exceeded 0.50 for all factors, supporting convergent validity. Most heterotrait-monotrait ratios were below 0.90, supporting discriminant validity. Cross-validation confirmed factorial replicability (CFI=0.917, TLI=0.904).
Conclusion
The TAILS-NS provides evidence of validity and reliability for measuring AI literacy among Thai nursing students and may serve as a standardized tool for curriculum evaluation. Future studies should expand the AI anxiety subscale and explore cross-cultural applicability in Southeast Asian nursing contexts.
Validity evaluation of a mock examination including multimedia-assisted items for the Korean Nursing Licensing Examination: a mixed-methods study with a counterbalanced design  
Sujin Shin, Inyoung Lee, Heejeong Kim, Rhayun Song, Jun-Ah Song, Hyun Sook Yi, Janghee Park, Mi Kyoung Yim, Minjae Lee, Subin Yu
J Educ Eval Health Prof. 2026;23:21.   Published online July 21, 2026
DOI: https://doi.org/10.3352/jeehp.2026.23.21    [Epub ahead of print]
  • 630 View
  • 106 Download
AbstractAbstract PDF
Purpose
This study aimed to evaluate the validity of multimedia-assisted items (MAIs) in a mock examination by examining their psychometric properties and nursing students’ perceptions.
Methods
This methodological study used a within-group counterbalanced design in which fourth-year nursing students in South Korea completed an online mock examination consisting of paired text-based items and MAIs. Psychometric properties were examined using classical test theory and item response theory, and learner perceptions and response processes were explored through post-examination surveys and focus group interviews.
Results
Among 515 participants, MAIs demonstrated overall difficulty and discrimination comparable to those of text-based items in both classical test theory and item response theory analyses, with similar test characteristic curves across counterbalanced groups, supporting measurement equivalence across item formats. Post-test surveys indicated that participants perceived MAIs as realistic reflections of clinical practice and as more helpful for problem solving and ability assessment than text-based items. Qualitative findings from focus group interviews (n=15) further supported the educational value of MAIs for assessing clinical judgment while emphasizing the need for appropriate multimedia use. An order effect was also observed: presenting MAIs first led to higher item response theory discrimination parameters in subsequent text-based items (P=0.020).
Conclusion
MAIs demonstrated psychometric properties comparable to those of text-based items and were positively perceived as valid tools for assessing clinical judgment and practical competence. These findings support the feasibility of incorporating multimedia-based assessment into the Korean Nursing Licensing Examination as part of a rigorous computer-based testing framework.
From accommodation to inclusion: understanding barriers and facilitators in clinical dental education for Chilean students with hearing disabilities: a qualitative study  
Josefa P. Retamal-Martínez
J Educ Eval Health Prof. 2026;23:20.   Published online June 29, 2026
DOI: https://doi.org/10.3352/jeehp.2026.23.20
  • 881 View
  • 108 Download
AbstractAbstract PDF
Purpose
This study aimed to describe inclusive mechanisms supporting dental students with hearing disabilities, focusing on facilitators of and barriers to participation in clinical education.
Methods
A qualitative case study was conducted at a Chilean university. Data were collected through semi-structured interviews with 2 students with hearing disabilities, 3 clinical instructors, and 1 Chilean Sign Language interpreter who works at the university and supports dental students with hearing disabilities, as well as through direct observation. Data were analyzed using inductive thematic analysis. Participant validation was used to enhance credibility.
Results
Facilitators were identified across pedagogical, technological, communicative, and relational domains. Pedagogical adaptations included early access to materials, visual resources, and flexible assessments. Technological tools such as real-time transcription supported communication and access to information. Alternative communication strategies included lip reading, written support, visual cues, and Chilean Sign Language. The presence of an interpreter was a key enabling factor. Relational support from faculty, peers, and family members also contributed substantially. Faculty flexibility and willingness to adapt were particularly important. Barriers were identified at structural, institutional, communicative, and attitudinal levels. Clinical environments presented physical and acoustic challenges, including noise, limited space, and complex interactions. Institutional responses were often reactive and involved limited planning. Communication barriers affected interactions with patients and peers, as well as academic literacy. Faculty reported limited training in inclusive education. Attitudinal barriers included misconceptions about disability and limited experience with disability.
Conclusion
Inclusive mechanisms are present but insufficiently systematized. Strengthening institutional planning and structural adaptations is essential to ensuring equitable clinical education.
Efficacy of a short multimodal workshop including stress management and communication skills training on objective structured clinical examination performance and exam-related anxiety in French medical students: a randomized controlled study
Meryl Vedrenne-Cloquet, Etienne de Montmollin, Augustin Gaudemer, Marion Strullu, Noémie De Cacqueray, Guillaume Beinse, Philibert Duriez, Marie-Aude Piot, Alain Cariou, Donia Bouzid, Damien Roux
J Educ Eval Health Prof. 2026;23:19.   Published online June 18, 2026
DOI: https://doi.org/10.3352/jeehp.2026.23.19
  • 893 View
  • 127 Download
AbstractAbstract PDF
Purpose
To assess the effectiveness of a short multimodal workshop, including stress-management and communication-skills training, on objective structured clinical examination (OSCE) performance and exam-related anxiety among medical students identified as needing additional support.
Methods
We conducted a prospective, open-label, randomized controlled trial including fifth-year medical students with the lowest scores on an initial mock OSCE. Participants were randomized 1:1 to the intervention group (stress management, communication training, and OSCE simulation) or the control group. OSCE examiners were blinded to group allocation. Anxiety was measured using the State-Trait Anxiety Inventory-State scale (STAI-State) before the intervention (T0), after the intervention (T1; intervention group only), and before the second mock OSCE (T2). OSCE scores and anxiety levels were compared between groups, and changes in anxiety over time were analyzed among students in the intervention group.
Results
A total of 109 students were included in the per-protocol analysis (49 intervention, 60 control). Global OSCE scores at T2 did not differ between groups (median [interquartile range], 13.0/20 [11.9–14.3] vs. 13.3/20 [12.0–14.6]; effect size, 0.08 [95% confidence interval, −0.14 to 0.29]; P=0.50). Anxiety levels were moderate, with no significant between-group difference (STAI-State score, 49 [43–50], n=8, in the intervention group vs. 44 [42–46], n=19, in the control group; P=0.21). A transient increase in situational anxiety was observed immediately after training in some intervention participants. No correlation was found between anxiety and OSCE performance.
Conclusion
A short multimodal workshop focused on stress management and communication training did not improve OSCE performance or OSCE-related anxiety, suggesting that single-session interventions may be insufficient.
Implementing patient- and family-centered rounds training with direct-observation assessment for medical students in the United States: a prospective feasibility study  
Emily Xiao, Erin Chen, Amit Pahwa
J Educ Eval Health Prof. 2026;23:18.   Published online June 18, 2026
DOI: https://doi.org/10.3352/jeehp.2026.23.18
  • 915 View
  • 77 Download
  • 1 Web of Science
  • 1 Crossref
AbstractAbstract PDFSupplementary Material
Purpose
This study evaluated the implementation of a patient- and family-centered rounds (PFCR) educational intervention and a standardized assessment form developed to assess changes in medical students’ PFCR performance over time.
Methods
From October 2023 to October 2024, medical students at Johns Hopkins School of Medicine attended a 1-hour PFCR simulation workshop during the Core Clerkship in Pediatrics. Students rotating at the main campus were assessed during rounds with a formative standardized form; all students received summative oral presentation and family rapport scores regardless of site. Performance was compared between students rotating in the first half (H1) and second half (H2) of the clerkship using Wilcoxon rank-sum tests with Holm correction for multiple comparisons. Linear mixed-effects models with student-level random intercepts were used to estimate changes across serial assessments.
Results
Among the 74 students rotating at the main campus, assessment forms were completed for 61% of students, with a median of 3 completed forms per student. H1 and H2 students had similar scores on both the formative assessment form and summative evaluations. Both groups improved significantly in inviting patient and family concerns across serial assessments. Summative scores did not differ between students evaluated before and after the educational intervention or between students at the main campus and those at community hospitals.
Conclusion
A structured PFCR educational session paired with a standardized assessment form was feasible to implement in a pediatric clerkship. Observed student-level median scores were at or above “meets expectations” across 6 PFCR domains, and no significant performance decline was observed among students rotating later in the clerkship. Further controlled studies are needed to determine whether this intervention improves PFCR skills compared with usual training.

Citations

Citations to this article as recorded by  
  • Two-Year Knowledge Retention After a Short Online Palliative Care Intervention for Medical Students: A Controlled Experimental Study
    Ozgur Tanriverdi, Kutluhan Cinbay
    Medical Science Educator.2026;[Epub]     CrossRef
Evaluation of a peer-assisted, simulation-based clinical skills training program in Spain: a prospective single-group before-and-after study  
Maria López-Brotons, Sergio Javaloy-Ballestero, José-Manuel Ramos-Rincón
J Educ Eval Health Prof. 2026;23:15.   Published online June 11, 2026
DOI: https://doi.org/10.3352/jeehp.2026.23.15
  • 789 View
  • 93 Download
AbstractAbstract PDFSupplementary Material
Purpose
This study aimed to evaluate the implementation and evolution of a peer-assisted, simulation-based clinical skills program at Miguel Hernández University, Spain, focusing on medical student participation, perceptions, and perceived educational value.
Methods
A prospective quasi-experimental pre–post study without a control group was conducted. Sessions were organized in small groups and led by senior student tutors under faculty supervision. Six workshops were offered during the first academic year and 8 during the second, with the addition of lumbar puncture and introductory clinical ultrasound. Knowledge acquisition was assessed using pre- and post-workshop questionnaires, and satisfaction was evaluated with Likert-scale surveys.
Results
A total of 139 unique students participated across the 2 academic years, with 77 participants in 2023–2024 and 77 in 2024–2025; 15 students participated in both academic years. These participants generated 440 workshop attendances. After incomplete questionnaires were excluded, 425 paired pre- and post-workshop evaluations were analyzed. Students reported improved learning, with workshop-specific knowledge test scores increasing significantly from 6.1±2.6 to 8.7±1.6 (Δ=2.5; 95% confidence interval [CI], 2.29–2.71; Cohen’s d=1.14; P<0.001). Male students showed greater knowledge gain than female students (Δ=2.8 vs. 2.3; 95% CI, 2.17–3.43 vs. 1.91–2.69; Cohen’s d=1.17 vs. 1.15; P=0.034), and second-year students improved more than third-year students (Δ=2.9 vs. 2.1; 95% CI, 2.36–3.44 vs. 1.67–2.53; Cohen’s d=1.21 vs. 1.11; P=0.001). Satisfaction was high, with mean scores above 4/5.
Conclusion
Clinical simulation combined with peer tutoring was feasible and well accepted in an undergraduate medical curriculum in Spain, achieving sustained participation over 2 academic years and consistently high satisfaction ratings. The program was associated with significant immediate improvements in workshop-specific knowledge test scores.
Reviews
Measurement properties of instruments assessing the clinical learning environment in health professions education: a COSMIN-based systematic review  
Matthias Michael Walter, Alexander P Schurz, Nele Adriaenssens, Evert Zinzen, Slavko Rogan
J Educ Eval Health Prof. 2026;23:14.   Published online June 11, 2026
DOI: https://doi.org/10.3352/jeehp.2026.23.14
  • 828 View
  • 88 Download
AbstractAbstract PDFSupplementary Material
Purpose
The clinical learning environment (CLE) is a crucial component of health professions education, providing the foundation for developing profession-specific clinical skills. This systematic review aimed to identify evaluated assessment tools for the CLE in health professions education and to report their measurement properties.
Methods
This systematic review was preregistered (IDESR000098), and its protocol was published previously. Eligible studies were peer-reviewed articles in English that developed and validated tools for assessing the CLE among undergraduate health professions students. The review followed the COSMIN guidelines for systematic reviews of patient-reported outcome measures. Multiple electronic databases, including MEDLINE, the Cochrane Library, ERIC, Education Research Complete, and CINAHL, were searched; studies were independently screened, and data were extracted. Data were synthesized using best-evidence synthesis according to COSMIN guidelines.
Results
Of the 6,236 records screened by title and abstract, 55 were eligible for full-text screening. A supplementary search identified 13 additional articles, resulting in 39 included articles. Overall, 28 tools were identified, with 4 tools (PET, PET-Midwifery, DECLEI, and MidSTEP) demonstrating sufficient content validity. Only MidSTEP demonstrated sufficient structural validity.
Conclusion
Only a minority of the included tools provided sufficient evidence of content and structural validity according to the COSMIN criteria. This finding indicates a systemic need for higher standards in monitoring clinical placements and identifies tools that should be re-evaluated and supported by additional research. Limitations include the exclusion of EMBASE and gray literature and the reliance on studies that predominantly used psychometric-first rather than content-validity-first designs.
Analysis of digital twin applications in nursing practice and education: a scoping review  
A Reum Lim, Hyun Kyoung Kim
J Educ Eval Health Prof. 2026;23:13.   Published online June 9, 2026
DOI: https://doi.org/10.3352/jeehp.2026.23.13
  • 1,309 View
  • 112 Download
  • 1 Web of Science
AbstractAbstract PDFSupplementary Material
This scoping review examined research applying digital twins in nursing practice and education and summarized their application domains, methods, outcomes, and implications. A human digital twin is a virtual health replica modeled from real-world data. This study followed the 5-stage scoping review process proposed by Arksey and O’Malley. Two researchers independently conducted the literature search without restrictions on publication year. From April 1 to 15, 2026, the Cochrane Library, PubMed, Embase, CINAHL, ERIC, and RISS databases were searched, and 15 studies were ultimately included. Digital twin applications were identified in 3 major domains: clinical practice and patient-centered care, education and training, and decision-making and workflow management. Application methods and outcomes varied according to technological implementation and included (1) modeling and data-driven prediction, (2) development of immersive learning and practice-training environments, and (3) system integration and decision-support frameworks. In clinical settings, multimodal patient data can be analyzed using artificial intelligence and machine learning to generate a virtual persona resembling the patient, thereby facilitating real-time personalized nursing care and self-management. In educational settings, digital twins can provide realistic and safe learning environments that enhance training effectiveness. Digital twins show substantial potential to advance predictive and personalized nursing in both clinical practice and education. Their data-driven capabilities are expected to contribute to innovative applications in future nursing practice and educational environments.
Research article
Development and initial validation of a brief 2-item measure of belonging among student physical therapists in the United States: a psychometric validation study  
Robyn Redline, Ashanti Jones, Thomas Gus Almonroeder
J Educ Eval Health Prof. 2026;23:11.   Published online May 28, 2026
DOI: https://doi.org/10.3352/jeehp.2026.23.11
  • 1,580 View
  • 93 Download
AbstractAbstract PDFSupplementary Material
Purpose
The objectives of this study were to develop a 2-item abbreviated version of the Program Sense of Belonging questionnaire (ProSBq) and to evaluate its ability to identify student physical therapists with relatively low valued competence and social acceptance.
Methods
A cross-sectional study was conducted using survey data from 634 students enrolled in physical therapist education programs across the United States. The 10-item ProSBq was used to assess 2 dimensions of belonging: valued competence and social acceptance. Principal component analysis was performed to identify representative items for each subscale, with 1 item selected per subscale. Pearson product-moment correlations were used to examine relationships between the single items and their corresponding subscale scores. Classification performance was evaluated by assessing how accurately the single-item responses classified students reporting a relatively low sense of valued competence and social acceptance, based on their full ProSBq subscale scores. Multiple single-item response thresholds were examined to assess classification accuracy.
Results
The single items demonstrated strong relationships with their corresponding subscale scores (r=0.63–0.80, with part-whole correction). For valued competence, sensitivity increased from 53.6% to 92.9%, whereas specificity decreased from 96.3% to 73.5% when a more inclusive threshold was used. A similar sensitivity-specificity tradeoff was observed for social acceptance. Receiver operating characteristic curve analyses demonstrated excellent discrimination (area under the curve ≥0.90).
Conclusion
Single ProSBq items demonstrated strong relationships with full valued competence and social acceptance subscale scores and acceptable classification performance. The abbreviated 2-item ProSBq may provide a practical and efficient method for identifying students experiencing low valued competence or social acceptance.
Brief report
Large language model-generated versus teacher-written objective structured clinical examination stations for medical students: a blinded comparative pilot study  
Piotr Szychowiak, Jonathan Wong-So, Hélène Messet, Mélanie Faure, Isaure Breteau, Simon Jamard, François Barbier, Maxime Desgrouas
J Educ Eval Health Prof. 2026;23:9.   Published online May 26, 2026
DOI: https://doi.org/10.3352/jeehp.2026.23.9
  • 1,215 View
  • 93 Download
AbstractAbstract PDFSupplementary Material
Developing objective structured clinical examination (OSCE) stations is time-consuming for medical teachers. We aimed to evaluate the ability of a large language model (LLM) to generate ready-to-use OSCE stations. Five OSCE stations generated by the LLM GPT-4o were evaluated by 7 expert assessors using a 5-point Likert scale and compared with 5 teacher-written stations targeting similar learning objectives. A station was considered to be of good quality if most assessors responded “agree” or “strongly agree” to the statement “The station is good enough to be used by students.” All teacher-written stations were rated as being of good quality, compared with only one GPT-4o-generated station. The LLM produced adequate clinical scenarios when reference knowledge was provided and tasks were clearly ordered, but it failed to generate reliable assessment grids. Careful review by teachers remained essential. GPT-4o failed to consistently produce fully ready-to-use OSCE stations.
Research articles
Developing clinical skills assessment modules for traditional, complementary, and integrative medicine in Korea: a participatory action research study  
Yoonjin Jeong, Seung Hwan Mun, Eunbyul Cho, Hye-Yoon Lee, Sang Woo Shin, Soyeon Kim, Eui-hyoung Hwang, Man-suk Hwang, Eunseok Kim, Jungyun Lee
J Educ Eval Health Prof. 2026;23:10.   Published online May 26, 2026
DOI: https://doi.org/10.3352/jeehp.2026.23.10
  • 1,466 View
  • 94 Download
AbstractAbstract PDFSupplementary Material
Purpose
This study aimed to develop pilot clinical skills assessment (CSA) modules for Korean medicine-specific procedures and to examine their preliminary appropriateness, perceived necessity, and feasibility as a foundation for future licensing-related assessment development.
Methods
A participatory action research framework, supplemented by qualitative interviews, was used to develop 4 CSA modules—acupuncture, Chuna manual therapy, pulse diagnosis, and constitutional diagnosis—in collaboration with expert evaluators, students, and standardized patients. The modules were implemented as formative examinations for third-year Korean medicine students, after which semi-structured interviews were conducted to obtain feedback on module content, implementation processes, and scoring procedures. Each module was also reviewed using the RUMBA checklist (Realistic, Understandable, Measurable, Behavioral, and Achievable), together with ratings of perceived necessity and feasibility for possible future use in licensing-related assessment. Interview data were analyzed inductively at the level of individual responses and then compared across modules and participant groups.
Results
Qualitative analysis yielded 3 themes: content and scoring criteria, physical environment or simulators, and education or training. Participants emphasized the need to make key aspects of performance more observable, improve authenticity through simulators or task trainers, and strengthen the capacity of scoring systems to distinguish between levels of student performance. Across all modules, mean RUMBA scores were high in the understandable, behavioral, and achievable domains, whereas measurability was more problematic, especially for pulse diagnosis.
Conclusion
These pilot findings clarify both the strengths and the limitations of Korean medicine-specific CSA modules. The modules received favorable ratings for understandability and achievability, whereas lower ratings for measurability and realism identified priorities for refinement before wider use. This study provides preliminary guidance for the continued development and broader evaluation of Korean medicine-specific performance assessments.
Personality type profiles of medical students and their differences by gender, age, and academic level in Korea: a cross-sectional study  
Yera Hur, Sanghee Yeo
J Educ Eval Health Prof. 2026;23:7.   Published online April 28, 2026
DOI: https://doi.org/10.3352/jeehp.2026.23.7
  • 1,708 View
  • 131 Download
AbstractAbstract PDFSupplementary Material
Purpose
Understanding the psychological characteristics of contemporary medical students is essential for effective educational design and learner support. This study aimed to identify medical students’ personality types using a geometric personality assessment tool (GEOPIA), determine whether differences exist by gender, age, or academic level, and explore the practical utility of such profiling for supporting educational practices in medical school settings.
Methods
The 40-item Korean Geometric Psychological Assessment (GEOPIA) was administered to 1,173 students across 5 Korean medical schools. GEOPIA classifies individuals into 4 primary types—Round (sociable, relationship-oriented), Triangle (task-oriented, challenging), Box (prudent, stability-seeking), and Curve (creative, sensitive). Frequency analyses and χ2 tests were conducted. Of the 1,016 respondents (response rate, 86.61%), 981 were included in the final analysis.
Results
The most common primary type was Round (40.3%), followed by Box (31.7%), Triangle (15.2%), and Curve (12.8%). Across the 12 combined profiles, Round–Box (21.9%) was the most prevalent, followed by Box–Round (19.0%) and Round–Triangle (9.7%). No significant differences were observed by gender (χ2=6.360, P=0.095, Cramer’s V=0.082), age (χ2=11.454, P=0.490, Cramer’s V=0.065), or academic level (χ2=18.044, P=0.260, Cramer’s V=0.078).
Conclusion
GEOPIA may provide a practical tool for identifying learner characteristics and supporting educational decision-making in medical school settings. In instructional design, personality-type data can inform group formation, activity planning, and assignment structure. In student support, the tool offers instructors and advisors a quick way to understand learners’ characteristics, which may help guide individualized counseling and promote effective learning experiences.
Review
Socioeconomic and economic factors affecting access and progression in medical schools: a systematic review and meta-analysis  
Arash Arianpoor, Alexia Pena, Annette Mercer, Jennifer Cox, Dimitra Lekkas, Francis Ruel Geronimo, Heidi Waldron, John Randal, Marcus Dabner, Marita Lynagh, Nalini Pather, Nicole Shepherd, Nigel Robb, Rose Berdin, Tim Wilkinson, Wendy Hu, Boaz Shulruf, Pin-Hsiang Huang
J Educ Eval Health Prof. 2026;23:6.   Published online April 16, 2026
DOI: https://doi.org/10.3352/jeehp.2026.23.6
  • 3,763 View
  • 150 Download
AbstractAbstract PDFSupplementary Material
Purpose
Socioeconomic disadvantage remains a major determinant of equitable access to, and progression within, medical education. This systematic review and meta-analysis examines both the impact and the magnitude of financial and economic disadvantage on student selection and progression in medical school.
Methods
Studies were included if they reported associations between socioeconomic indicators (e.g., parental income, occupation, education, geographic deprivation, or premedical debt) and selection or progression outcomes, and were excluded if they lacked clearly defined economic predictors or sufficient data for binary effect sizes. Searches were conducted across PubMed, Scopus, ERIC, Embase, ProQuest, and EBSCO (2005–2025). Study selection employed an active machine-learning screening process. Extracted data included sample characteristics, socioeconomic measures, and outcome types, with risk of bias assessed using the Risk of Bias Instrument. Random-effects meta-analysis was conducted where appropriate.
Results
Thirty-two studies of medical programs were included, yielding 28 effect sizes for selection and 9 for progression. Household economic and educational disadvantage, identified through parental indices, was consistently associated with reduced odds of admission (odds ratio [OR], 0.6; 95% confidence interval [CI], 0.55–0.65) and poorer progression (OR, 0.56; 95% CI, 0.53–0.59). Geographic deprivation also exerted a negative effect, particularly on selection (OR, 0.69; 95% CI, 0.5–0.93).
Conclusion
Socioeconomic disadvantage exerts a pervasive influence across the medical education continuum. Addressing these inequities requires sustained financial, academic, and psychosocial support both before and during their studies. Students’ economic circumstances should therefore be considered in medical school selection policy and curriculum development to further enhance equity within medical schools and the profession.
Research article
Implementing low-cost 3D-printed brain coloring activities in neuroanatomy teaching for medical students in Singapore: a cross-sectional study  
Jason Wen Yau Lee, Fernando Bello, Jai Prashanth Rao
J Educ Eval Health Prof. 2026;23:5.   Published online March 18, 2026
DOI: https://doi.org/10.3352/jeehp.2026.23.5
  • 2,256 View
  • 215 Download
  • 1 Web of Science
  • 1 Crossref
AbstractAbstract PDFSupplementary Material
Purpose
Three-dimensional (3D)–printed models have been increasingly used in medical education, but most studies have focused on satisfaction or outcomes following isolated learning activities. This study aimed to explore students’ perceptions of learning, engagement, usability, and learning strategies after completing a series of neuroanatomy-related coloring activities using a low-cost 3D-printed model.
Methods
This cross-sectional study involved Year 1 medical students at Duke-NUS Medical School. Students participated in 3 structured coloring activities using a modular 3D-printed brain model during a neuroanatomy session. An anonymous survey was administered 1 week after the third activity to assess students’ perceived learning value, engagement (behavioral, cognitive, emotional, and agentic), usability, and learning strategies using Likert-scale items and open-ended questions.
Results
A total of 48 students completed the survey, and the instrument showed acceptable to high internal consistency. Students reported high perceived learning value, positive engagement across multiple domains during the coloring activity, and high usability of the model. Participation in the learning activities was associated with significantly higher behavioral and agentic engagement, perceived learning value, and greater use of learning strategies than non-participation. Overall, active manipulation and hands-on exploration were perceived as beneficial for learning.
Conclusion
Low-cost 3D-printed brain models may serve as valuable learning tools to complement existing anatomy teaching approaches when paired with well-designed learning activities. Students reported positive learning experiences and high engagement during the activities. These findings highlight the importance of sound pedagogical design and curriculum integration to maximize learning.

Citations

Citations to this article as recorded by  
  • Neuroanatomy in medical education: Trends and perspectives from a bibliometric and thematic analysis
    Piotr Wysocki, Aleksandra Wyciślok, Andrzej Żytkowski
    Translational Research in Anatomy.2026; 45: 100513.     CrossRef
Review
Performance of large language models in medical licensing examinations: a systematic review and meta-analysis  
Haniyeh Nouri, Abdollah Mahdavi, Ali Abedi, Alireza Mohammadnia, Mahnaz Hamedan, Masoud Amanzadeh
J Educ Eval Health Prof. 2025;22:36.   Published online November 18, 2025
DOI: https://doi.org/10.3352/jeehp.2025.22.36
  • 4,919 View
  • 258 Download
  • 5 Web of Science
  • 10 Crossref
AbstractAbstract PDFSupplementary Material
Purpose
This study systematically evaluates and compares the performance of large language models (LLMs) in answering medical licensing examination questions. By conducting subgroup analyses based on language, question format, and model type, this meta-analysis aims to provide a comprehensive overview of LLM capabilities in medical education and clinical decision-making.
Methods
This systematic review, registered in PROSPERO and following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines, searched MEDLINE (PubMed), Scopus, and Web of Science for relevant articles published up to February 1, 2025. The search strategy included Medical Subject Headings (MeSH) terms and keywords related to (“ChatGPT” OR “GPT” OR “LLM variants”) AND (“medical licensing exam*” OR “medical exam*” OR “medical education” OR “radiology exam*”). Eligible studies evaluated LLM accuracy on medical licensing examination questions. Pooled accuracy was estimated using a random-effects model, with subgroup analyses by LLM type, language, and question format. Publication bias was assessed using Egger’s regression test.
Results
This systematic review identified 2,404 studies. After removing duplicates and excluding irrelevant articles through title and abstract screening, 36 studies were included after full-text review. The pooled accuracy was 72% (95% confidence interval, 70.0% to 75.0%) with high heterogeneity (I2=99%, P<0.001). Among LLMs, GPT-4 achieved the highest accuracy (81%), followed by Bing (79%), Claude (74%), Gemini/Bard (70%), and GPT-3.5 (60%) (P=0.001). Performance differences across languages (range, 62% in Polish to 77% in German) were not statistically significant (P=0.170).
Conclusion
LLMs, particularly GPT-4, can match or exceed medical students’ examination performance and may serve as supportive educational tools. However, due to variability and the risk of errors, they should be used cautiously as complements rather than replacements for traditional learning methods.

Citations

Citations to this article as recorded by  
  • Effective prompt design for large language models in clinical practice
    Steven Callens
    Acta Clinica Belgica.2026; 81(2): 118.     CrossRef
  • Examination of Gemini's ability to answer anatomical questions: An overview
    D. Chytas, G. Noussios, M.-K. Kaseta, D. Chrysikos, A.V. Vasiliadis, C. Lyrtzis, T. Troupis
    Morphologie.2026; 110(369): 101118.     CrossRef
  • Evaluación comparativa de modelos de inteligencia artificial de última generación frente a psiquiatras humanos en el examen nacional de subespecialidad en Perú: un estudio transversal
    Javier A. Flores-Cohaila, Jeff Huarcaya-Victoria, Cesar Copaja-Corzo
    Educación Médica.2026; 27(3): 101179.     CrossRef
  • ChatGPT vs Claude: Scoping Review with ☸️SAIMSARA

    SAIMSARA Journal.2026;[Epub]     CrossRef
  • PeruMedQA: A Stress Evaluation Using Ten Large Language Models to Answer Medical Exams
    Rodrigo M. Carrillo-Larco
    Medical Science Educator.2026; 36(3): 1091.     CrossRef
  • Evaluating the accuracy and communication quality of large language models in Ewing sarcoma: a comparative analysis of ChatGPT, Claude, Gemini, DeepSeek, and Grok
    Cihan Ünyılmaz
    Frontiers in Pediatrics.2026;[Epub]     CrossRef
  • The Promises and Perils of Clinical Decision Support Artificial Intelligence
    Jorge Cervantes, Bhavya Vashi
    The Clinical Teacher.2026;[Epub]     CrossRef
  • NASA TLX workload profiles of multimodal artificial intelligence models and dental students during objective structured practical dental examinations
    Sanaa N. Al-Haj Ali, Ra’fat I. Farah
    Discover Education.2026;[Epub]     CrossRef
  • Artificial Intelligence and Sleep
    Logan Douglas Schneider, John Hernandez, Conor Heneghan
    Neurologic Clinics.2026;[Epub]     CrossRef
  • Generative artificial intelligence in medical education: from knowledge assessment to clinical reasoning and professional competence
    Renxian Xie, Beien Zhang, Lifeng Xiao
    Frontiers in Medicine.2026;[Epub]     CrossRef
Research articles
Perceptions of faculty and medical students regarding an undergraduate research culture activity in Myanmar: a qualitative study  
Htain Lin Aung, Moe Oo Thant, July Maung Maung, Ye Hlaing Oo, Thin Thin Toe, Hla Moe
J Educ Eval Health Prof. 2025;22:33.   Published online October 27, 2025
DOI: https://doi.org/10.3352/jeehp.2025.22.33
  • 2,753 View
  • 287 Download
AbstractAbstract PDFSupplementary Material
Purpose
This study explored the perceptions of faculty members and third-year medical students regarding the research culture activity (RCA), a program designed to engage undergraduates in research at the University of Medicine, Mandalay, Myanmar. It aimed to identify the knowledge, attitudes, and skills (KAS) gained, the challenges encountered, and suggestions for improvement.
Methods
This qualitative study employed 4 semi-structured focus group discussions with 17 third-year medical students and 16 faculty members who participated in the 2020 RCA. Student responses related to KAS were analyzed using a deductive framework approach, while challenges and suggestions were examined through inductive thematic analysis. Discussions were audio-recorded, transcribed verbatim in Burmese, translated into English, and collaboratively coded using Atlas.ti version 9.0.5.
Results
Participants reported improved understanding of scientific literature, greater responsibility, strengthened teamwork, and enhanced practical research skills. Reported challenges included limited research preparedness, scheduling conflicts, inconsistent supervision, financial constraints, and weak coordination with inpatient clinicians. Participants also suggested clearer guidelines, pre-research training, protected time, stronger supervision, and institutional budgetary support.
Conclusion
The RCA provides substantial educational value in developing research competencies and remains a promising, potentially adaptable model for resource-limited settings. Its sustainability will depend on institutional commitment, supervisory capacity, and modest financial investment. Future research should prospectively assess KAS outcomes, compare supervision models and group sizes, evaluate digital workflows for efficiency, and conduct long-term follow-up of graduates’ scholarly activities to build evidence for scalable implementation.
Development and psychometric assessment of a scale for evaluating healthcare professionals’ attitudes toward interprofessional education and collaboration in the United States: a cross-sectional study  
Michael Christopher Banks, Ryan Brock Mutcheson, Maedot Ariaya Haymete, Serkan Toy
J Educ Eval Health Prof. 2025;22:32.   Published online October 20, 2025
DOI: https://doi.org/10.3352/jeehp.2025.22.32
  • 2,532 View
  • 261 Download
AbstractAbstract PDFSupplementary Material
Purpose
Interprofessional education (IPE) is increasingly recognized as critical to preparing health professionals for collaborative practice, yet rigorous assessment remains limited by a lack of psychometrically sound instruments. Building on a previously developed questionnaire for physicians, this study aimed to expand the scale to include allied health professionals and to evaluate whether the factor structure remained consistent across professions. We hypothesized that a similar factor structure would emerge from the combined dataset, thereby supporting the scale’s generalizability.
Methods
This observational study included 930 healthcare professionals in the United States (379 physicians, 419 nurses, 76 pharmacists, and others) who completed a 35-item questionnaire addressing IPE competency domains. Data were collected between December 2019 and May 2020. Exploratory factor analysis was employed to examine the factor structure, followed by item response theory (IRT) analyses to assess item fit, reliability, and validity. Raw data are available upon request.
Results
Factor analysis of 22 retained items confirmed a 5-factor solution: teamwork and communication, patient-centered care, roles and responsibilities, ethics and attitudes, and reflective practice, explaining 59% of the variance. Subscale reliabilities ranged from α=0.65 to 0.87. IRT analyses supported construct validity and measurement precision, while identifying areas for refinement in reflective practice.
Conclusion
This study demonstrates that the scale is reliable, valid, and generalizable across diverse health professions. It provides a robust tool for assessing attitudes toward IPE, offering value for curriculum evaluation, institutional benchmarking, and future longitudinal research on professional identity formation and collaborative practice.
The impact of differential item functioning on ability estimation using the Korean Medical Licensing Examination with computerized adaptive testing: a post-hoc simulation study  
Dogyeong Kim, Jeongwook Choi, Dong Gi Seo
J Educ Eval Health Prof. 2025;22:31.   Published online October 10, 2025
DOI: https://doi.org/10.3352/jeehp.2025.22.31
  • 3,805 View
  • 200 Download
AbstractAbstract PDFSupplementary Material
Purpose
This study examined the impact of differential item functioning (DIF) on ability estimation in a computerized adaptive testing (CAT) environment using real response data from the 2017 Korean Medical Licensing Examination (KMLE). We hypothesized that excluding gender-based DIF items would improve estimation accuracy, particularly for examinees at the extremes of the ability scale.
Methods
The study was conducted in 2 steps: (1) DIF detection and (2) post-hoc simulation. The analysis used data from 3,259 examinees who completed all 360 dichotomous items. Gender-based DIF was detected with the residual-based DIF method (reference group: males; focal group: females). Two CAT conditions (all items vs. DIF-excluded) were compared against a “true θ” estimated from a fixed-form test of 264 non-DIF items. Accuracy was evaluated using bias, root mean square error (RMSE), and correlation with true θ.
Results
In the CAT condition excluding DIF items, accuracy improved, with RMSE reduced and correlation with true θ increased. However, bias was slightly larger in magnitude. Gender-specific analyses showed that DIF removal reduced the underestimation of female ability but increased the underestimation of male ability, yielding estimates that were fairer across genders. When DIF items were included, estimation errors were more pronounced at both low and high ability levels.
Conclusion
Managing DIF in CAT-based high-stakes examinations can enhance fairness and precision. Using real examinee data, this study provides practical evidence of the implications of DIF for CAT-based measurement and supports fairness-oriented test design.
Performance of GPT-4o and o1-Pro on United Kingdom Medical Licensing Assessment-style items: a comparative study  
Behrad Vakili, Aadam Ahmad, Mahsa Zolfaghari
J Educ Eval Health Prof. 2025;22:30.   Published online October 10, 2025
DOI: https://doi.org/10.3352/jeehp.2025.22.30
  • 3,207 View
  • 254 Download
AbstractAbstract PDFSupplementary Material
Purpose
Large language models (LLMs) such as ChatGPT, and their potential to support autonomous learning for licensing exams like the UK Medical Licensing Assessment (UKMLA), are of growing interest. However, empirical evaluations of artificial intelligence (AI) performance against the UKMLA standard remain limited.
Methods
We evaluated the performance of 2 recent ChatGPT versions, GPT-4o and o1-Pro, on a curated set of 374 UKMLA-style single-best-answer items spanning diverse medical specialties. Statistical comparisons using McNemar’s test assessed the significance of differences between the 2 models. Specialties were analyzed to identify domain-specific variation. In addition, 20 image-based items were evaluated.
Results
GPT-4o achieved an accuracy of 88.8%, while o1-Pro achieved 93.0%. McNemar’s test revealed a statistically significant difference in favor of o1-Pro. Across specialties, both models demonstrated excellent performance in surgery, psychiatry, and infectious diseases. Notable differences arose in dermatology, respiratory medicine, and imaging, where o1-Pro consistently outperformed GPT-4o. Nevertheless, isolated weaknesses in general practice were observed. The analysis of image-based items showed 75% accuracy for GPT-4o and 90% for o1-Pro (P=0.25).
Conclusion
ChatGPT shows strong potential as an adjunct learning tool for UKMLA preparation, with both models achieving scores above the calculated pass mark. This underscores the promise of advanced AI models in medical education. However, specialty-specific inconsistencies suggest AI tools should complement, rather than replace, traditional study methods.
Leveraging feedback mechanisms to improve the quality of objective structured clinical examinations in Singapore: an exploratory action research study  
Han Ting Jillian Yeo, Dujeepa Dasharatha Samarasekera, Michael Dean
J Educ Eval Health Prof. 2025;22:28.   Published online September 30, 2025
DOI: https://doi.org/10.3352/jeehp.2025.22.28
  • 4,052 View
  • 190 Download
AbstractAbstract PDFSupplementary Material
Purpose
Variability in examiner scoring threatens the fairness and reliability of objective structured clinical examinations (OSCEs). While examiner standardization exists, there is currently no structured, psychometric-informed, individualized feedback mechanism for examiners. This study explored the feasibility and perceived value of such a mechanism using an action research approach to co-design and iteratively refine examiner feedback reports.
Methods
Two exploratory cycles were conducted between November 2023 and June 2024 with phase 4 OSCE examiners at the Yong Loo Lin School of Medicine. In cycle 1, psychometric analyses of examiner scoring for a phase 4 OSCE informed the design of individualized reports, which were evaluated through interviews. Revisions were made to the format of the report and implemented in cycle 2, where examiner responses were again collected. Data were analyzed thematically, supported by reflective logs and field notes.
Results
Nine examiners participated in cycle 1 and 7 in cycle 2. In cycle 1, examiners highlighted challenges in interpreting complex terminology, leading to report refinements such as glossaries and visual graphs. In cycle 2, examiners demonstrated greater confidence in applying feedback, requested longitudinal reports, and shifted from initial resistance to reflective engagement. Across cycles, the reports improved credibility, neutrality, and examiner self-regulation.
Conclusion
This exploratory study suggests that psychometric-informed feedback reports can facilitate examiner reflection and transparency in OSCEs. While the findings highlight feasibility and examiner acceptance, longitudinal delivery of feedback, collection of quantitative outcome data, and larger samples are needed to establish whether such reports improve scoring consistency and assessment fairness.
Performance of ChatGPT-4 on the French Board of Plastic Reconstructive and Aesthetic Surgery written exam: a descriptive study
Emma Dejean-Bouyer, Anoujat Kanlagna, François Thuau, Pierre Perrot, Ugo Lancien
J Educ Eval Health Prof. 2025;22:27.   Published online September 30, 2025
DOI: https://doi.org/10.3352/jeehp.2025.22.27
  • 2,048 View
  • 176 Download
AbstractAbstract PDFSupplementary Material
Purpose
This study aims to evaluate the performance of Chat Generative Pre-Trained Transformer 4 (ChatGPT-4) on the French Board of Plastic, Reconstructive, and Aesthetic Surgery written examination and to assess its role as a supplementary resource in helping residents prepare for the qualification examination in plastic surgery.
Methods
This descriptive study evaluated ChatGPT-4’s performance on 213 items from the October 2024 French Board of Plastic, Reconstructive, and Aesthetic Surgery written examination. Responses were assessed for accuracy, logical reasoning, internal and external information use, and were categorized for fallacies by independent reviewers. Statistical analyses included chi-square tests and Fisher’s exact test for significance.
Results
ChatGPT-4 answered all questions across the 10 modules, achieving an overall accuracy rate of 77.5%. The model applied logical reasoning in 98.1% of the questions, utilized internal information in 94.4%, and incorporated external information in 91.1%.
Conclusion
ChatGPT-4 performs satisfactorily on the French Board of Plastic, Reconstructive, and Aesthetic Surgery written examination. Its accuracy met the minimum passing standards for the exam. While responses generally align with expected knowledge, careful verification remains necessary, particularly for questions involving image interpretation. As artificial intelligence continues to evolve, ChatGPT-4 is expected to become an increasingly reliable tool for medical education. At present, it remains a valuable resource for assisting plastic surgery residents in their training.
Validity of the formative physical therapy Student and Clinical Instructor Performance Assessment Instrument in the United States: a quasi-experimental, time-series study  
Sean Gallivan, Jamie Bayliss
J Educ Eval Health Prof. 2025;22:26.   Published online September 26, 2025
DOI: https://doi.org/10.3352/jeehp.2025.22.26
  • 2,206 View
  • 185 Download
AbstractAbstract PDFSupplementary Material
Purpose
The aim of this study was to assess the validity of the Student and Clinical Instructor Performance Instrument (SCIPAI), a novel formative tool used in physical therapist education to assess student and clinical instructor (CI) performance throughout clinical education experiences (CEEs). The researchers hypothesized that the SCIPAI would demonstrate concurrent, predictive, and construct validity while offering additional contemporary validity evidence.
Methods
This quasi-experimental, time-series study had 811 student-CI pairs complete 2 SCIPAIs before after CEE midpoint, and an endpoint Clinical Performance Instrument (CPI) during beginning to terminal CEEs in a 1-year period. Spearman rank correlation analyses used final SCIPAI and CPI like-item scores to assess concurrent validity; and earlier SCIPAI and final CPI like-item scores to assess predictive validity. Construct validity was assessed via progression of student and CI performance scores within CEEs using Wilcoxon signed-rank testing. No randomization/grouping of subjects occurred.
Results
Moderate correlation existed between like final SCIPAI and CPI items (P<0.005) and between some like items of earlier SCIPAIs and final CPIs (P<0.005). Student performance scores demonstrated progress from SCIPAIs 1 to 4 within CEEs (P<0.005). While a greater number of CIs demonstrated progression rather than regression in performance from SCIPAI 1 to SCIPAI 4, the greater magnitude of decreases in CI performance contributed to an aggregate ratings decrease of CI performance (P<0.005).
Conclusion
The SCIPAI demonstrates concurrent, predictive, and construct validity when used by students and CIs to rate student performance at regular points throughout clinical education experiences.
Proposal for setting a passing score for the Korean Nursing Licensing Examination
Janghee Park, Mi Kyoung Yim, Sujin Shin, Rhayun Song, Jun-Ah Song, Inyoung Lee, Heejeong Kim, Minjae Lee
J Educ Eval Health Prof. 2025;22:25.   Published online September 8, 2025
DOI: https://doi.org/10.3352/jeehp.2025.22.25
  • 2,830 View
  • 246 Download
AbstractAbstract PDFSupplementary Material
Purpose
The Korean Nursing Licensing Examination (KNLE) is planning to transition to a computer-based test (CBT). This study aims to propose a reasonable and efficient method for setting passing scores.
Methods
A standard-setting (passing score setting) analysis was conducted using an expert panel over the past 3 years of the national nursing examination. The standard-setting method was modified from Angoff, and the validity of the passing score was verified through the Hofstee method. The standard-setting workshop was conducted in 2 stages: first, a pilot workshop for 2 subjects, followed by a second workshop where 6 additional subjects were selected based on the pilot results. For items with an actual correct answer rate of 90% or higher, the estimated correct answer rate for minimum competency was calculated using the observed correct answer rate. A survey and discussion with the expert panel were also conducted regarding the standard-setting procedures and results.
Results
The passing score for the national nursing examination was calculated using the new method, and the score was slightly higher than the existing score. The nursing subject had similar results; however, the legal subjects varied.
Conclusion
The modified Angoff and Hofstee methods were successfully applied to the KNLE. Using the actual correct answer rate as an indicator to derive expected minimum competency was shown to be effective. This approach could streamline future standard-setting processes, particularly when converting to CBT.
Decline in attrition rates in United States pediatric residency and fellowship programs, 2007–2020: a repeated cross-sectional study  
Emma Omoruyi, Greg Russell, Kimberly Montez
J Educ Eval Health Prof. 2025;22:24.   Published online September 5, 2025
DOI: https://doi.org/10.3352/jeehp.2025.22.24
  • 2,480 View
  • 171 Download
AbstractAbstract PDFSupplementary Material
Purpose
Declining fill rates in US pediatric residency and subspecialty programs requires trainee retention. Attrition, defined as transfers, withdrawals, dismissals, unsuccessful completions, or deaths, disrupts program function and impacts the pediatric workforce pipeline. It aims to evaluate attrition trends among pediatric residents and fellows in Accreditation Council for Graduate Medical Education (ACGME)-accredited programs from 2007 to 2020.
Methods
This repeated cross-sectional study analyzed publicly available ACGME Data Resource Book records. Attrition rates and 95% confidence intervals (CIs) were calculated overall and by subspecialty. Logistic regression assessed temporal changes; odds ratios (ORs) compared 2020 to 2007.
Results
From 2007–2020, pediatric residents increased from 8,145 to 9,419 and fellows from 2,875 to 4,279. Aggregate annual resident attrition averaged 1.71% (range, 0.93%–2.64%), and fellow attrition ranged from 12.39%–30.87%. Transfer rates declined from 18.05 to 5.20 per 1,000 trainees (P<0.0001), withdrawals from 5.65 to 2.76 (P=0.030), and dismissals from 3.14 in 2010 to 1.27 in 2020 (P=0.0068). Odds of unsuccessful completion significantly decreased in categorical pediatrics (OR, 0.41; 95% CI, 0.29–0.58), pediatric cardiology (OR, 0.08; 95% CI, 0.01–0.64), pediatric critical care (OR, 0.14; 95% CI, 0.06–0.35), and neonatal-perinatal medicine (OR, 0.46; 95% CI, 0.20–1.08).
Conclusion
Although attrition has improved, premature trainee loss can still disrupt program operations and threaten workforce development. Attrition may reflect educational environment quality, support structures, or selection processes. Greater data transparency is needed to understand demographic trends and inform equitable retention strategies, ultimately strengthening training programs and sustaining the United States pediatric workforce.
Comparing generative artificial intelligence platforms and nursing student performance on a women’s health nursing examination in Korea: a Rasch model approach  
Eun Jeong Ko, Tae Kyung Lee, Geum Hee Jeong
J Educ Eval Health Prof. 2025;22:23.   Published online September 5, 2025
DOI: https://doi.org/10.3352/jeehp.2025.22.23
  • 2,735 View
  • 243 Download
  • 2 Web of Science
  • 1 Crossref
AbstractAbstract PDFSupplementary Material
Purpose
This psychometric study aimed to compare the ability parameter estimates of generative artificial intelligence (AI) platforms with those of nursing students on a 50-item women’s health nursing examination at Hallym University, Korea, using the Rasch model. It also sought to estimate item difficulty parameters and evaluate AI performance across varying difficulty levels.
Methods
The exam, consisting of 39 multiple-choice items and 11 true/false items, was administered to 111 fourth-year nursing students in June 2023. In December 2024, 6 generative AI platforms (GPT-4o, ChatGPT free version, Claude.ai, Clova X, Mistral.ai, Google Gemini) completed the same items. The responses were analyzed using the Rasch model to estimate the ability and difficulty parameters. Unidimensionality was verified by the Dimensionality Evaluation to Enumerate Contributing Traits (DETECT), and analyses were conducted using the R packages irtQ and TAM.
Results
The items satisfied unidimensionality (DETECT=–0.16). Item difficulty parameter estimates ranged from –3.87 to 1.96 logits (mean=–0.61), with a mean difficulty index of 0.79. Examinees’ ability parameter estimates ranged from –0.71 to 3.15 logits (mean=1.17). GPT-4o, ChatGPT free version, and Claude.ai outperformed the median student ability (1.09 logits), scoring 2.68, 2.34, and 2.34, respectively, while Clova X, Mistral.ai, and Google Gemini exhibited lower scores (0.20, –0.12, 0.80). The test information curve peaked below θ=0, indicating suitability for examinees with low to average ability.
Conclusion
Advanced generative AI platforms approximated the performance of high-performing students, but outcomes varied. The Rasch model effectively evaluated AI competency, supporting its potential utility for future AI performance assessments in nursing education.

Citations

Citations to this article as recorded by  
  • Eroding scholarly integrity: Confronting the misuse of generative AI in nursing education
    Kechi Iheduru-Anderson
    Nurse Education Today.2026; 164: 107145.     CrossRef
Brief Report
Effectiveness of interprofessional education enhanced by live consultation observations for healthcare students and new professionals in Singapore: a retrospective cross-sectional study  
Lynette Mei Lim Goh, Wai Leong Chiu, Sky Wei Chee Koh
J Educ Eval Health Prof. 2025;22:21.   Published online August 21, 2025
DOI: https://doi.org/10.3352/jeehp.2025.22.21
  • 4,187 View
  • 196 Download
AbstractAbstract Supplementary Material
This study aims to evaluate whether incorporating live consultation observations into interprofessional education (IPE) improves learning evaluation scores among healthcare professionals and students. A retrospective cross-sectional analysis was conducted using evaluation data from AHP IPE sessions held from January 2020 to December 2023 across 7 primary care clinics in Singapore. Evaluation scores were compared between sessions with facilitated discussions only (n=667) and sessions with additional live consultation observations (n=501). Logistic regression was used to analyze factors associated with perfect evaluation scores. Sessions that included live consultations were significantly more likely to achieve perfect evaluation scores (odds ratio [OR], 1.68; 95% confidence interval [CI], 1.27–2.22). Nursing/care coordinator and allied health professions (OR 2.07 and 1.76 respectively) were significantly more likely to give perfect scores compared to medical professions. Healthcare professionals were also more likely to give perfect scores than students (OR, 1.52; 95% CI,1.08–2.14), indicating enhanced perceived effectiveness. These findings support the use of experiential learning strategies to optimize interprofessional training outcomes.
Research articles
Comparison between GPT-4 and human raters in grading pharmacy students’ exam responses in Malaysia: a cross-sectional study
Wuan Shuen Yap, Pui San Saw, Li Ling Yeap, Shaun Wen Huey Lee, Wei Jin Wong, Ronald Fook Seng Lee
J Educ Eval Health Prof. 2025;22:20.   Published online July 28, 2025
DOI: https://doi.org/10.3352/jeehp.2025.22.20
  • 5,157 View
  • 321 Download
  • 1 Web of Science
AbstractAbstract PDFSupplementary Material
Purpose
Manual grading is time-consuming and prone to inconsistencies, prompting the exploration of generative artificial intelligence tools such as GPT-4 to enhance efficiency and reliability. This study investigated GPT-4’s potential in grading pharmacy students’ exam responses, focusing on the impact of optimized prompts. Specifically, it evaluated the alignment between GPT-4 and human raters, assessed GPT-4’s consistency over time, and determined its error rates in grading pharmacy students’ exam responses.
Methods
We conducted a comparative study using past exam responses graded by university-trained raters and by GPT-4. Responses were randomized before evaluation by GPT-4, accessed via a Plus account between April and September 2024. Prompt optimization was performed on 16 responses, followed by evaluation of 3 prompt delivery methods. We then applied the optimized approach across 4 item types. Intraclass correlation coefficients and error analyses were used to assess consistency and agreement between GPT-4 and human ratings.
Results
GPT-4’s ratings aligned reasonably well with human raters, demonstrating moderate to excellent reliability (intraclass correlation coefficient=0.617–0.933), depending on item type and the optimized prompt. When stratified by grade bands, GPT-4 was less consistent in marking high-scoring responses (Z=–5.71–4.62, P<0.001). Overall, despite achieving substantial alignment with human raters in many cases, discrepancies across item types and a tendency to commit basic errors necessitate continued educator involvement to ensure grading accuracy.
Conclusion
With optimized prompts, GPT-4 shows promise as a supportive tool for grading pharmacy students’ exam responses, particularly for objective tasks. However, its limitations—including errors and variability in grading high-scoring responses—require ongoing human oversight. Future research should explore advanced generative artificial intelligence models and broader assessment formats to further enhance grading reliability.
Longitudinal relationships between Korean medical students’ academic performance in medical knowledge and clinical performance examinations: a retrospective longitudinal study  
Yulim Kang, Hae Won Kim
J Educ Eval Health Prof. 2025;22:18.   Published online June 10, 2025
DOI: https://doi.org/10.3352/jeehp.2025.22.18
  • 5,491 View
  • 291 Download
AbstractAbstract PDFSupplementary Material
Purpose
This study investigated the longitudinal relationships between performance on 3 examinations assessing medical knowledge and clinical skills among Korean medical students in the clinical phase. This study addressed the stability of each examination score and the interrelationships among examinations over time.
Methods
A retrospective longitudinal study was conducted at Yonsei University College of Medicine in Korea with a cohort of 112 medical students over 2 years. The students were in their third year in 2022 and progressed to the fourth year in 2023. We obtained comprehensive clinical science examination (CCSE) and progress test (PT) scores 3 times (T1–T3), and clinical performance examination (CPX) scores twice (T1 and T2). Autoregressive cross-lagged models were fitted to analyze their relationships.
Results
For each of the 3 examinations, the score at 1 time point predicted the subsequent score. Regarding cross-lagged effects, the CCSE at T1 predicted PT at T2 (β=0.472, P<0.001) and CCSE at T2 predicted PT at T3 (β=0.527, P<0.001). The CPX at T1 predicted the CCSE at T2 (β=0.163, P=0.006), and the CPX at T2 predicted the CCSE at T3 (β=0.154, P=0.006). The PT at T1 predicted the CPX at T2 (β=0.273, P=0.006).
Conclusion
The study identified each examination’s stability and the complexity of the longitudinal relationships between them. These findings may help predict medical students’ performance on subsequent examinations, potentially informing the provision of necessary student support.
Educational/Faculty development material
Radiotorax.es: a web-based tool for formative self-assessment in chest X-ray interpretation  
Verónica Illescas-Megías, Jorge Manuel Maqueda-Pérez, Dolores Domínguez-Pinos, Teodoro Rudolphi Solero, Francisco Sendra-Portero
J Educ Eval Health Prof. 2025;22:17.   Published online June 9, 2025
DOI: https://doi.org/10.3352/jeehp.2025.22.17
  • 9,259 View
  • 2,147,483,870 Download
AbstractAbstract PDFSupplementary Material
Radiotorax.es is a free, non-profit web-based tool designed to support formative self-assessment in chest X-ray interpretation. This article presents its structure, educational applications, and usage data from 11 years of continuous operation. Users complete interpretation rounds of 20 clinical cases, compare their reports with expert evaluations, and conduct a structured self-assessment. From 2011 to 2022, 14,389 users registered, and 7,726 completed at least one session. Most were medical students (75.8%), followed by residents (15.2%) and practicing physicians (9.0%). The platform has been integrated into undergraduate medical curricula and used in various educational contexts, including tutorials, peer and expert review, and longitudinal tracking. Its flexible design supports self-directed learning, instructor-guided use, and multicenter research. As a freely accessible resource based on real clinical cases, Radiotorax.es provides a scalable, realistic, and well-received training environment that promotes diagnostic skill development, reflection, and educational innovation in radiology education.
Research articles
Performance of large language models on Thailand’s national medical licensing examination: a cross-sectional study  
Prut Saowaprut, Romen Samuel Wabina, Junwei Yang, Lertboon Siriwat
J Educ Eval Health Prof. 2025;22:16.   Published online May 12, 2025
DOI: https://doi.org/10.3352/jeehp.2025.22.16
  • 7,625 View
  • 376 Download
  • 6 Web of Science
  • 5 Crossref
AbstractAbstract PDFSupplementary Material
Purpose
This study aimed to evaluate the feasibility of general-purpose large language models (LLMs) in addressing inequities in medical licensure exam preparation for Thailand’s National Medical Licensing Examination (ThaiNLE), which currently lacks standardized public study materials.
Methods
We assessed 4 multi-modal LLMs (GPT-4, Claude 3 Opus, Gemini 1.0/1.5 Pro) using a 304-question ThaiNLE Step 1 mock examination (10.2% image-based), applying deterministic API configurations and 5 inference repetitions per model. Performance was measured via micro- and macro-accuracy metrics compared against historical passing thresholds.
Results
All models exceeded passing scores, with GPT-4 achieving the highest accuracy (88.9%; 95% confidence interval, 88.7–89.1), surpassing Thailand’s national average by more than 2 standard deviations. Claude 3.5 Sonnet (80.1%) and Gemini 1.5 Pro (72.8%) followed hierarchically. Models demonstrated robustness across 17 of 20 medical domains, but variability was noted in genetics (74.0%) and cardiovascular topics (58.3%). While models demonstrated proficiency with images (Gemini 1.0 Pro: +9.9% vs. text), text-only accuracy remained superior (GPT-4o: 90.0% vs. 82.6%).
Conclusion
General-purpose LLMs show promise as equitable preparatory tools for ThaiNLE Step 1. However, domain-specific knowledge gaps and inconsistent multi-modal integration warrant refinement before clinical deployment.

Citations

Citations to this article as recorded by  
  • Is artificial intelligence getting better at anatomy? A two‐year review of ChatGPT's free public versions
    Bahattin Paslı, Ceren Günenç Beşer
    Anatomical Sciences Education.2026; 19(8): 1279.     CrossRef
  • The performance of ChatGPT and other large language models on multiple‐choice questions in biomedical disciplines: A meta‐analysis
    Colleen M. Cheverko, Volodymyr Mavrych, Olena Bolgova, Fathima Raahima Riyas Mohamed, Jennifer Westrick, Lorena Juarez, Emily Rush, Kathryn A. Solka, Alison F. Doubleday, Jessica N. Byram, Robert Becker, Victoria Gomez, Brenda K. Anak Ganeng, Leslie A. Ho
    Anatomical Sciences Education.2026; 19(9): 1556.     CrossRef
  • Performance of GPT-4o and o1-Pro on United Kingdom Medical Licensing Assessment-style items: a comparative study
    Behrad Vakili, Aadam Ahmad, Mahsa Zolfaghari
    Journal of Educational Evaluation for Health Professions.2025; 22: 30.     CrossRef
  • Large Language Models for the National Radiological Technologist Licensure Examination in Japan: Cross-Sectional Comparative Benchmarking and Evaluation of Model-Generated Items Study
    Toshimune Ito, Toru Ishibashi, Tatsuya Hayashi, Shinya Kojima, Kazumi Sogabe
    JMIR Medical Education.2025; 11: e81807.     CrossRef
  • Technologies, opportunities, challenges, and future directions for integrating generative artificial intelligence into medical education: a narrative review
    Junseok Kang, Jihyun Ahn
    Ewha Medical Journal.2025; 48(4): e53.     CrossRef
Evaluation of a virtual objective structured clinical examination in the metaverse (Second Life) to assess the clinical skills in emergency radiology of medical students in Spain: a cross-sectional study  
Alba Virtudes Perez-Baena, Teodoro Rudolphi-Solero, Rocio Lorenzo-Alvarez, Dolores Dominguez-Pinos, Miguel Jose Ruiz-Gomez, Francisco Sendra-Portero
J Educ Eval Health Prof. 2025;22:12.   Published online April 21, 2025
DOI: https://doi.org/10.3352/jeehp.2025.22.12
  • 6,826 View
  • 2,147,483,947 Download
  • 3 Web of Science
  • 5 Crossref
AbstractAbstract PDFSupplementary Material
Purpose
The objective structured clinical examination (OSCE) is an effective but resource-intensive tool for assessing clinical competence. This study hypothesized that implementing a virtual OSCE in the Second Life (SL) platform in the metaverse as a cost-effective alternative will effectively assess and enhance clinical skills in emergency radiology while being feasible and well-received. The aim was to evaluate a virtual radiology OSCE in SL as a formative assessment, focusing on feasibility, educational impact, and students’ perceptions.
Methods
Two virtual 6-station OSCE rooms dedicated to emergency radiology were developed in SL. Sixth-year medical students completed the OSCE during a 1-hour session in 2022–2023, followed by feedback including a correction checklist, individual scores, and group comparisons. Students completed a questionnaire with Likert-scale questions, a 10-point rating, and open-ended comments. Quantitative data were analyzed using the Student t-test and the Mann-Whitney U test, and qualitative data through thematic analysis.
Results
In total, 163 students participated, achieving mean scores of 5.1±1.4 and 4.9±1.3 (out of 10) in the 2 virtual OSCE rooms, respectively (P=0.287). One hundred seventeen students evaluated the OSCE, praising the teaching staff (9.3±1.0), project organization (8.8±1.2), OSCE environment (8.7±1.5), training usefulness (8.6±1.5), and formative self-assessment (8.5±1.4). Likert-scale questions and students’ open-ended comments highlighted the virtual environment’s attractiveness, case selection, self-evaluation usefulness, project excellence, and training impact. Technical difficulties were reported by 13 students (8%).
Conclusion
This study demonstrated the feasibility of incorporating formative OSCEs in SL as a useful teaching tool for undergraduate radiology education, which was cost-effective and highly valued by students.

Citations

Citations to this article as recorded by  
  • Effectiveness of VR and traditional training in medical education for mass casualty management: an OSCE-based randomized controlled trial
    Zhe Li, Wan Chen, Guozheng Qiu, Lei Shi, Yutao Tang, Xibin Xu, Sanshan Zhu, Liwen Lyu
    BMC Medical Education.2026;[Epub]     CrossRef
  • ECOE virtual de radiología en el metaverso Second Life®: comparación de estudiantes de tercer y sexto curso
    Francisco Sendra-Portero, Alba Virtudes Pérez-Baena, Teodoro Rudolphi-Solero, Rocío Lorenzo-Álvarez, Dolores Domínguez-Pinos, Miguel José Ruiz-Gómez
    Educación Médica.2026; 27(3): 101167.     CrossRef
  • A roadmap to implement Metaverse in education field based on ADDIE model: A systematic review
    Yu Shi, Wei Wei Goh, N.Z. Jhanjhi
    Computers and Education Open.2026; 10: 100365.     CrossRef
  • Metaverse-based objective structured clinical examinations: an exploratory approach to advancing clinical competency assessment
    Yeon-Ju Huh, Joon Sung Shin, Narae Yoon, Ju Whi Kim, Do Hoon Kim, Chanwoong Kim, Seoi Jeong, Yejin Yoon, Soyeon Shin, Hyoun-Joong Kong, Sun Jung Myung
    Korean Journal of Medical Education.2026; 38(2): 139.     CrossRef
  • An AI‐Driven H5P Training for Epilepsy Emergency Management Preparedness in Dental Education
    İlkem Gökçe, Özlem Sürel Karabilgin Öztürkçü, Ozan Karaca, Cihan Varol, Mukadder İnci Başer Kolcu
    European Journal of Dental Education.2026;[Epub]     CrossRef
A nationwide survey on the curriculum and educational resources related to the Clinical Skills Test of the Korean Medical Licensing Examination: a cross-sectional descriptive study  
Eun-Kyung Chung, Seok Hoon Kang, Do-Hoon Kim, MinJeong Kim, Ji-Hyun Seo, Keunmi Lee, Eui-Ryoung Han
J Educ Eval Health Prof. 2025;22:11.   Published online March 13, 2025
DOI: https://doi.org/10.3352/jeehp.2025.22.11
  • 7,091 View
  • 351 Download
  • 1 Web of Science
  • 1 Crossref
AbstractAbstract PDFSupplementary Material
Purpose
The revised Clinical Skills Test (CST) of the Korean Medical Licensing Exam aims to provide a better assessment of physicians’ clinical competence and ability to interact with patients. This study examined the impact of the revised CST on medical education curricula and resources nationwide, while also identifying areas for improvement within the revised CST.
Methods
This study surveyed faculty responsible for clinical clerkships at 40 medical schools throughout Korea to evaluate the status and changes in clinical skills education, assessment, and resources related to the CST. The researchers distributed the survey via email through regional consortia between December 7, 2023 and January 19, 2024.
Results
Nearly all schools implemented preliminary student–patient encounters during core clinical rotations. Schools primarily conducted clinical skills assessments in the third and fourth years, with a simplified form introduced in the first and second years. Remedial education was conducted through various methods, including one-on-one feedback from faculty after the assessment. All schools established clinical skills centers and made ongoing improvements. Faculty members did not perceive the CST revisions as significantly altering clinical clerkship or skills assessments. They suggested several improvements, including assessing patient records to improve accuracy and increasing the objectivity of standardized patient assessments to ensure fairness.
Conclusion
During the CST, students’ involvement in patient encounters and clinical skills education increased, improving the assessment and feedback processes for clinical skills within the curriculum. To enhance students’ clinical competencies and readiness, strengthening the validity and reliability of the CST is essential.

Citations

Citations to this article as recorded by  
  • Nationwide cross-sectional survey on the necessity of including a clinical skills assessment in the national licensure examination for Doctors of Korean Medicine
    Aram Jeong, Eunbyul Cho, Chan-Young Kwon, Sanghoon Lee, Chungsik Cho, Sangwoo Shin, Min Hwangbo, Dong-Hyeon Kim, Hye-Yoon Lee
    Medicine.2025; 104(45): e45366.     CrossRef
Simulation-based teaching versus traditional small group teaching for first-year medical students among high and low scorers in respiratory physiology, India: a randomized controlled trial  
Nalini Yelahanka Channegowda, Dinker Ramanand Pai, Shivasakthy Manivasakan
J Educ Eval Health Prof. 2025;22:8.   Published online February 21, 2025
DOI: https://doi.org/10.3352/jeehp.2025.22.8
  • 4,708 View
  • 376 Download
  • 2 Web of Science
  • 4 Crossref
AbstractAbstract PDFSupplementary Material
Purpose
Although it is widely utilized in clinical subjects for skill training, using simulation-based education (SBE) for teaching basic science concepts to phase I medical students or pre-clinical students is limited. Simulation-based education/teaching is preferred in cardiovascular and respiratory physiology when compared to other systems because it is easy to recreate both the normal physiological component and alterations in the simulated environment, thus a promoting deep understanding of the core concepts.
Methods
A block randomized study was conducted among 107 phase 1 (first-year) medical undergraduate students at a Deemed to be University in India. Group A received SBE and Group B traditional small group teaching. The effectiveness of the teaching intervention was assessed using pre- and post-tests. Student feedback was obtained through a self administered structured questionnaire via an anonymous online survey and by in-depth interview.
Results
The intervention group showed a statistically significant improvement in post-test scores compared to the control group. A sub-analysis revealed that high scorers performed better than low scorers in both groups, but the knowledge gain among low scorers was more significant in the intervention group.
Conclusion
This teaching strategy offers a valuable supplement to traditional methods, fostering a deeper comprehension of clinical concepts from the outset of medical training.

Citations

Citations to this article as recorded by  
  • Use of simulation in physiology teaching: a focus on medical education
    Levi Davis, Madeline Rapp, Zachary Rundell, Layla Al-Nakkash
    Current Opinion in Physiology.2026; 47: 100920.     CrossRef
  • Jigsaw Cooperative Learning: Reflections on Topic‐Contingent Preference
    Zhanglei Mu
    The Clinical Teacher.2026;[Epub]     CrossRef
  • Simulation and Augmented Reality on Academic Performance and Engagement in Grade 11 Earth and Life Science
    Abigail G. Dumaguing, Wilfred G. Alava Jr.
    International Journal of Innovative Science and Research Technology.2025; : 2817.     CrossRef
  • Education, Neuroscience, and Technology: A Review of Applied Models
    Elena Granado De la Cruz, Francisco Javier Gago-Valiente, Óscar Gavín-Chocano, Eufrasio Pérez-Navío
    Information.2025; 16(8): 664.     CrossRef
Reliability and construct validation of the Blended Learning Usability Evaluation–Questionnaire with interprofessional clinicians in Canada: a methodological study  
Anish Kumar Arora, Jeff Myers, Tavis Apramian, Kulamakan Kulasegaram, Daryl Bainbridge, Hsien Seow
J Educ Eval Health Prof. 2025;22:5.   Published online January 16, 2025
DOI: https://doi.org/10.3352/jeehp.2025.22.5
  • 6,321 View
  • 361 Download
  • 1 Web of Science
  • 4 Crossref
AbstractAbstract PDFSupplementary Material
Purpose
To generate Cronbach’s alpha and further mixed methods construct validity evidence for the Blended Learning Usability Evaluation–Questionnaire (BLUE-Q).
Methods
Forty interprofessional clinicians completed the BLUE-Q after finishing a 3-month long blended learning professional development program in Ontario, Canada. Reliability was assessed with Cronbach’s α for each of the 3 sections of the BLUE-Q and for all quantitative items together. Construct validity was evaluated through the Grand-Guillaume-Perrenoud et al. framework, which consists of 3 elements: congruence, convergence, and credibility. To compare quantitative and qualitative results, descriptive statistics, including means and standard deviations for each Likert scale item of the BLUE-Q were calculated.
Results
Cronbach’s α was 0.95 for the pedagogical usability section, 0.85 for the synchronous modality section, 0.93 for the asynchronous modality section, and 0.96 for all quantitative items together. Mean ratings (with standard deviations) were 4.77 (0.506) for pedagogy, 4.64 (0.654) for synchronous learning, and 4.75 (0.536) for asynchronous learning. Of the 239 qualitative comments received, 178 were identified as substantive, of which 88% were considered congruent and 79% were considered convergent with the high means. Among all congruent responses, 69% were considered confirming statements and 31% were considered clarifying statements, suggesting appropriate credibility. Analysis of the clarifying statements assisted in identifying 5 categories of suggestions for program improvement.
Conclusion
The BLUE-Q demonstrates high reliability and appropriate construct validity in the context of a blended learning program with interprofessional clinicians, making it a valuable tool for comprehensive program evaluation, quality improvement, and evaluative research in health professions education.

Citations

Citations to this article as recorded by  
  • Multi-methods development and validation of a tool for use in measuring serious illness communication competence: Assessment of clinical encounters – Communication tool (ACE-CT)
    Anish K. Arora, Hsien Seow, Daryl Bainbridge, Kulamakan Kulasegaram, Tavis Apramian, Nadia Incardona, Leah Steinberg, Justin Sanders, Zhimeng Jia, Oren Levine, Jessica Simon, Karen Zhang, Zelda Freitas, Clare Fuller, Amanda Lee Roze des Ordons, Jill Dombr
    Patient Education and Counseling.2026; 144: 109465.     CrossRef
  • A Mixed Methods Framework for Usability Profiling in Learning Environments: Integrating Adapted Heuristics and Machine Learning
    Wilintonn Fidel Ortiz-Fajardo, Andrés Felipe Solis-Pino, Cesar Collazos, Gustavo E. Constain
    TecnoLógicas.2026;[Epub]     CrossRef
  • Utilizing cognitive interview in the item refinement of the Blended Teaching Assessment Tool (BTAT) for Health Professions Education
    Maria Teresita B. Dalusong, Glenda Sanggalang Ogerio, Valentin C. Dones, Maria Elizabeth M. Grageda
    Philippine Journal of Health Research and Development.2025; 29(2): 54.     CrossRef
  • All providers Better Communication Skills (ABCs) program: protocol for a randomized controlled trial assessing communication training effectiveness with interprofessional clinicians
    Hsien Seow, Anish K. Arora, Daryl Bainbridge, Zhimeng Jia, Leah Steinberg, Nadia Incardona, Oren Levine, Justin J. Sanders, Jessica Simon, Amanda Roze des Ordons, Karen Zhang, Jeff Myers
    BMC Palliative Care.2025;[Epub]     CrossRef
Empathy and tolerance of ambiguity in medical students and doctors participating in art-based observational training at the Rijksmuseum in Amsterdam, the Netherlands: a before-and-after study  
Stella Anna Bult, Thomas van Gulik
J Educ Eval Health Prof. 2025;22:3.   Published online January 14, 2025
DOI: https://doi.org/10.3352/jeehp.2025.22.3
  • 7,058 View
  • 375 Download
  • 5 Web of Science
  • 5 Crossref
AbstractAbstract PDFSupplementary Material
Purpose
This research presents an experimental study using validated questionnaires to quantitatively assess the outcomes of art-based observational training in medical students, residents, and specialists. The study tested the hypothesis that art-based observational training would lead to measurable effects on judgement skills (tolerance of ambiguity) and empathy in medical students and doctors.
Methods
An experimental cohort study with pre- and post-intervention assessments was conducted using validated questionnaires and qualitative evaluation forms to examine the outcomes of art-based observational training in medical students and doctors. Between December 2023 and June 2024, 15 art courses were conducted in the Rijksmuseum in Amsterdam. Participants were assessed on empathy using the Jefferson Scale of Empathy (JSE) and tolerance of ambiguity using the Tolerance of Ambiguity in Medical Students and Doctors (TAMSAD) scale.
Results
In total, 91 participants were included; 29 participants completed the JSE and 62 completed the TAMSAD scales. The results showed statistically significant post-test increases for mean JSE and TAMSAD scores (3.71 points for the JSE, ranging from 20 to 140, and 1.86 points for the TAMSAD, ranging from 0 to 100). The qualitative findings were predominantly positive.
Conclusion
The results suggest that incorporating art-based observational training in medical education improves empathy and tolerance of ambiguity. This study highlights the importance of art-based observational training in medical education in the professional development of medical students and doctors.

Citations

Citations to this article as recorded by  
  • Understanding uncertainty and ambiguity in medicine and medical education: a narrative review with implications for training
    Sarine Sarkis, Christian Raphael
    Postgraduate Medical Journal.2026; 102(1207): 461.     CrossRef
  • Observational training for surgical residents using visual arts in the museum
    Thomas M. van Gulik, Stella A. Bult, Pien E.J. de Ruiter, Floortje Huizing, Alexander de Mol van Otterloo, Alexander Leijdesdorff, Sjoerd Lagarde
    Surgery.2026; 190: 109843.     CrossRef
  • Training the eye and diagnosing the canvas in the Museum ‘A perspective on art-based medical education’
    T.M. van Gulik, S.A. Bult, P.E.J. de Ruiter, F. Huizing, A. Leijdesdorff, S. Lagarde, A. de Mol van Otterloo
    Ethics, Medicine and Public Health.2026; 34: 101243.     CrossRef
  • Developing a Feasible Arts and Humanities Course Using Visual Thinking Strategies and Haiku Writing: A Mixed-Methods Study
    Hirohisa Fujikawa, Takayuki Ando, Junji Haruta
    Medical Science Educator.2025; 35(6): 3105.     CrossRef
  • Erb’s Palsy: Visual Diagnosis in Art before Medical History?
    Pien E.J. de Ruiter, Stella A. Bult, Jeroen R. Dijkstra, Thomas M. van Gulik
    Gynecologic and Obstetric Investigation.2025; 91(1): 26.     CrossRef
Inter-rater reliability and content validity of the measurement tool for portfolio assessments used in the Introduction to Clinical Medicine course at Ewha Womans University College of Medicine: a methodological study  
Dong-Mi Yoo, Jae Jin Han
J Educ Eval Health Prof. 2024;21:39.   Published online December 10, 2024
DOI: https://doi.org/10.3352/jeehp.2024.21.39
  • 6,591 View
  • 291 Download
  • 4 Web of Science
  • 5 Crossref
AbstractAbstract PDFSupplementary Material
Purpose
This study aimed to examine the reliability and validity of a measurement tool for portfolio assessments in medical education. Specifically, it investigated scoring consistency among raters and assessment criteria appropriateness according to an expert panel.
Methods
A cross-sectional observational study was conducted from September to December 2018 for the Introduction to Clinical Medicine course at the Ewha Womans University College of Medicine. Data were collected for 5 randomly selected portfolios scored by a gold-standard rater and 6 trained raters. An expert panel assessed the validity of 12 assessment items using the content validity index (CVI). Statistical analysis included Pearson correlation coefficients for rater alignment, the intraclass correlation coefficient (ICC) for inter-rater reliability, and the CVI for item-level validity.
Results
Rater 1 had the highest Pearson correlation (0.8916) with the gold-standard rater, while Rater 5 had the lowest (0.4203). The ICC for all raters was 0.3821, improving to 0.4415 after excluding Raters 1 and 5, indicating a 15.6% reliability increase. All assessment items met the CVI threshold of ≥0.75, with some achieving a perfect score (CVI=1.0). However, items like “sources” and “level and degree of performance” showed lower validity (CVI=0.72).
Conclusion
The present measurement tool for portfolio assessments demonstrated moderate reliability and strong validity, supporting its use as a credible tool. For a more reliable portfolio assessment, more faculty training is needed.

Citations

Citations to this article as recorded by  
  • Values Education in Curriculum Reform: A Qualitative Document Analysis of the Türkiye Century Maarif Model for Primary Education
    Ethem Gürhan
    Journal of Educational Research and Practice.2026; 4(1): 24.     CrossRef
  • Developing the assessment tool for TVET institutional accreditation: a modified Delphi method
    Sagar Mani Neupane, Bhanu Pandit, Pramod Bhakta Acharya, Prakash C. Bhattarai
    Frontiers in Education.2026;[Epub]     CrossRef
  • Large Language Models as Clinical Support Tools in Drug Information Services: Performance Comparison With Pharmacists
    Nuntapong Boonrit, Najwa Bin‐useng, Aphichaya Sirijariyawat, Tanawan Chaithong, Suppakun Rattayarungsri, Sutima Yodsri, Ashley M. Hopkins, Warit Ruanglertboon
    JACCP: JOURNAL OF THE AMERICAN COLLEGE OF CLINICAL PHARMACY.2026;[Epub]     CrossRef
  • Monocular Markerless Motion Capture for Concept-Stage Human-Factors Evaluation of an IVD Sample-Loading Unit
    Ming Guo, Mingfeng He, Chencan Wang, Qingyun Liu, Shenyan Ma, Yuhan Li
    Applied Sciences.2026; 16(16): 8086.     CrossRef
  • On the quantitative analysis of assessment scores with implicit and explicit constraints
    Sanjeeb Shrestha, Xiaoying Kong, Paul Kwan
    Studies in Educational Evaluation.2025; 87: 101509.     CrossRef
Validation of the 21st Century Skills Assessment Scale for public health students in Thailand: a methodological study  
Suphawadee Panthumas, Kaung Zaw, Wirin Kittipichai
J Educ Eval Health Prof. 2024;21:37.   Published online December 10, 2024
DOI: https://doi.org/10.3352/jeehp.2024.21.37
  • 7,348 View
  • 498 Download
  • 1 Crossref
AbstractAbstract PDFSupplementary Material
Purpose
This study aimed to develop and validate the 21st Century Skills Assessment Scale (21CSAS) for Thai public health (PH) undergraduate students using the Partnership for 21st Century Skills framework.
Methods
A cross-sectional survey was conducted among 727 first- to fourth-year PH undergraduate students from 4 autonomous universities in Thailand. Data were collected using self-administered questionnaires between January and March 2023. Exploratory factor analysis (EFA) was used to explore the underlying dimensions of 21CSAS, while confirmatory factor analysis (CFA) was conducted to test the hypothesized factor structure using Mplus software (Muthén & Muthén). Reliability and item discrimination were assessed using Cronbach’s α and the corrected item-total correlation, respectively.
Results
EFA performed on a dataset of 300 students revealed a 20-item scale with a 6-factor structure: (1) creativity and innovation; (2) critical thinking and problem-solving; (3) information, media, and technology; (4) communication and collaboration; (5) initiative and self-direction; and (6) social and cross-cultural skills. The rotated eigenvalues ranged from 2.12 to 1.73. CFA performed on another dataset of 427 students confirmed a good model fit (χ2/degrees of freedom=2.67, comparative fit index=0.93, Tucker-Lewis index=0.91, root mean square error of approximation=0.06, standardized root mean square residual=0.06), explaining 34%–71% of variance in the items. Item loadings ranged from 0.58 to 0.84. The 21CSAS had a Cronbach’s α of 0.92.
Conclusion
The 21CSAS proved be a valid and reliable tool for assessing 21st century skills among Thai PH undergraduate students. These findings provide insights for educational system to inform policy, practice, and research regarding 21st-century skills among undergraduate students.

Citations

Citations to this article as recorded by  
  • Generative AI use and sustainable pedagogical practices as predictors of 4Cs skill development: The mediating role of digital literacy among international postgraduate students in China
    Muhammad Azram, Hong Mei, Asim Khan, Bilal Ahmad, Uneeb Ur Rehman Ali
    Information Processing & Management.2027; 64(2): 105119.     CrossRef
History article
History of the medical licensure system in Korea from the late 1800s to 1992
Sang-Ik Hwang
J Educ Eval Health Prof. 2024;21:36.   Published online December 9, 2024
DOI: https://doi.org/10.3352/jeehp.2024.21.36
  • 8,506 View
  • 140 Download
  • 1 Web of Science
  • 2 Crossref
AbstractAbstract PDFSupplementary Material
The introduction of modern Western medicine in the late 19th century, notably through vaccination initiatives, marked the beginning of governmental involvement in medical licensure, with the licensing of doctors who performed vaccinations. The establishment of the national medical school “Euihakkyo” in 1899 further formalized medical education and licensure, granting graduates the privilege to practice medicine without additional examinations. The enactment of the Regulations on Doctors in 1900 by the Joseon government aimed to define doctor qualifications, including modern and traditional practitioners, comprehensively. However, resistance from the traditional medical community hindered its full implementation. During the Japanese colonial occupation of the Korean Peninsula from 1910 to 1945, the medical licensure system was controlled by colonial authorities, leading to the marginalization of traditional Korean medicine and the imposition of imperial hierarchical structures. Following liberation in 1945 from Japanese colonial rule, the Korean government undertook significant reforms, culminating in the National Medical Law, which was enacted in 1951. This law redefined doctor qualifications and reinstated the status of traditional Korean medicine. The introduction of national examinations for physicians increased state involvement in ensuring medical competence. The privatization of the Korean Medical Licensing Examination led to the establishment of the Korea Health Personnel Licensing Examination Institute in 1992, which assumed responsibility for administering licensing examinations for all healthcare workers. This shift reflected a move towards specialized management of professional standards. The evolution of the medical licensure system in Korea illustrates a dynamic process shaped by the historical context, balancing the protection of public health with the rights of medical practitioners.

Citations

Citations to this article as recorded by  
  • Public servants or privileged elites? -Analyzing physician strikes in the Republic of Korea
    Seongwon Choi
    Social Science & Medicine.2026; 395: 119092.     CrossRef
  • Educational Innovation in Nursing Management after Reforms in the National Nurse Licensure Examination: A Narrative Review
    Yoomi Jung
    Journal of Korean Academy of Nursing Administration.2026; 32(3): 137.     CrossRef
Research article
Effectiveness of ChatGPT-4o in developing continuing professional development plans for graduate radiographers: a descriptive study  
Minh Chau, Elio Stefan Arruzza, Kelly Spuur
J Educ Eval Health Prof. 2024;21:34.   Published online November 18, 2024
DOI: https://doi.org/10.3352/jeehp.2024.21.34
  • 5,899 View
  • 267 Download
  • 6 Web of Science
  • 8 Crossref
AbstractAbstract PDFSupplementary Material
Purpose
This study evaluates the use of ChatGPT-4o in creating tailored continuing professional development (CPD) plans for radiography students, addressing the challenge of aligning CPD with Medical Radiation Practice Board of Australia (MRPBA) requirements. We hypothesized that ChatGPT-4o could support students in CPD planning while meeting regulatory standards.
Methods
A descriptive, experimental design was used to generate 3 unique CPD plans using ChatGPT-4o, each tailored to hypothetical graduate radiographers in varied clinical settings. Each plan followed MRPBA guidelines, focusing on computed tomography specialization by the second year. Three MRPBA-registered academics assessed the plans using criteria of appropriateness, timeliness, relevance, reflection, and completeness from October 2024 to November 2024. Ratings underwent analysis using the Friedman test and intraclass correlation coefficient (ICC) to measure consistency among evaluators.
Results
ChatGPT-4o generated CPD plans generally adhered to regulatory standards across scenarios. The Friedman test indicated no significant differences among raters (P=0.420, 0.761, and 0.807 for each scenario), suggesting consistent scores within scenarios. However, ICC values were low (–0.96, 0.41, and 0.058 for scenarios 1, 2, and 3), revealing variability among raters, particularly in timeliness and completeness criteria, suggesting limitations in the ChatGPT-4o’s ability to address individualized and context-specific needs.
Conclusion
ChatGPT-4o demonstrates the potential to ease the cognitive demands of CPD planning, offering structured support in CPD development. However, human oversight remains essential to ensure plans are contextually relevant and deeply reflective. Future research should focus on enhancing artificial intelligence’s personalization for CPD evaluation, highlighting ChatGPT-4o’s potential and limitations as a tool in professional education.

Citations

Citations to this article as recorded by  
  • Shaping the Future of Radiography Education: Lessons From ChatGPT and Generative AI
    Minh T. Chau, Haydn Kerr, Clare L. Singh, Bismark Ofori‐Manteaw, Elio Arruzza, Kelly Bentley‐Spuur
    Journal of Medical Radiation Sciences.2026; 73(3): 327.     CrossRef
  • Applied insights for using Generative Artificial Intelligence in Faculty Development in Health Professions Education
    Melchor Sánchez-Mendiola, Megan Anakin, Ardi Findyartini, Rachel Levine, Ana Da Silva, Farhan Saeed Vakani
    MedEdPublish.2026; 15: 279.     CrossRef
  • Halted medical education and medical residents’ training in Korea, journal metrics, and appreciation to reviewers and volunteers
    Sun Huh
    Journal of Educational Evaluation for Health Professions.2025; 22: 1.     CrossRef
  • The ‘Negotiator’: Assessing artificial intelligence (AI) interview preparation for graduate radiographers
    M. Chau, E. Arruzza, C.L. Singh
    Journal of Medical Imaging and Radiation Sciences.2025; 56(5): 101982.     CrossRef
  • ‘Bill’: An artificial intelligence (AI) clinical scenario coach for medical radiation science education
    M. Chau, G. Higgins, E. Arruzza, C.L. Singh
    Radiography.2025; 31(5): 103002.     CrossRef
  • Exploring ChatGPT-4o-generated reflections: Alignment with professional standards in diagnostic radiography: A pilot experiment
    C Nabasenja, M Chau, E Green
    Journal of Medical Imaging and Radiation Sciences.2025; 56(6): 102082.     CrossRef
  • Applied insights for using Generative Artificial Intelligence in Faculty Development in Health Professions Education
    Melchor Sánchez-Mendiola, Megan Anakin, Ardi Findyartini, Rachel Levine, Ana Da Silva, Farhan Saeed Vakani
    MedEdPublish.2025; 15: 279.     CrossRef
  • A research roadmap for AI opportunities in student assessment for medical education
    Morteza Rezaei-Zadeh, Magdalena Cerbin-Koczorowska
    BMC Medical Education.2025;[Epub]     CrossRef
Technical report
Increased accessibility of computer-based testing for residency application to a hospital in Brazil with item characteristics comparable to paper-based testing: a psychometric study  
Marcos Carvalho Borges, Luciane Loures Santos, Paulo Henrique Manso, Elaine Christine Dantas Moisés, Pedro Soler Coltro, Priscilla Costa Fonseca, Paulo Roberto Alves Gentil, Rodrigo de Carvalho Santana, Lucas Faria Rodrigues, Benedito Carlos Maciel, Hilton Marcos Alves Ricz
J Educ Eval Health Prof. 2024;21:32.   Published online November 11, 2024
DOI: https://doi.org/10.3352/jeehp.2024.21.32
  • 3,352 View
  • 192 Download
  • 1 Crossref
AbstractAbstract PDFSupplementary Material
Purpose
With the coronavirus disease 2019 pandemic, online high-stakes exams have become a viable alternative. This study evaluated the feasibility of computer-based testing (CBT) for medical residency applications in Brazil and its impacts on item quality and applicants’ access compared to paper-based testing.
Methods
In 2020, an online CBT was conducted in a Ribeirao Preto Clinical Hospital in Brazil. In total, 120 multiple-choice question items were constructed. Two years later, the exam was performed as paper-based testing. Item construction processes were similar for both exams. Difficulty and discrimination indexes, point-biserial coefficient, difficulty, discrimination, guessing parameters, and Cronbach’s α coefficient were measured based on the item response and classical test theories. Internet stability for applicants was monitored.
Results
In 2020, 4,846 individuals (57.1% female, mean age of 26.64±3.37 years) applied to the residency program, versus 2,196 individuals (55.2% female, mean age of 26.47±3.20 years) in 2022. For CBT, there was an increase of 2,650 applicants (120.7%), albeit with significant differences in demographic characteristics. There was a significant increase in applicants from more distant and lower-income Brazilian regions, such as the North (5.6% vs. 2.7%) and Northeast (16.9% vs. 9.0%). No significant differences were found in difficulty and discrimination indexes, point-biserial coefficients, and Cronbach’s α coefficients between the 2 exams.
Conclusion
Online CBT with multiple-choice questions was a viable format for a residency application exam, improving accessibility without compromising exam integrity and quality.

Citations

Citations to this article as recorded by  
  • Chatbot Underperformance in Biology and Image-Based Questions in Medical Education
    Joyce Santana Rizzi, Lorraine Silva Requena, Angelica Maria Bicudo, Pedro Tadao Hamamoto Filho, Renato Ferretti
    Journal of CME.2025;[Epub]     CrossRef
Research articles
Validation of the Blended Learning Usability Evaluation–Questionnaire (BLUE-Q) through an innovative Bayesian questionnaire validation approach  
Anish Kumar Arora, Charo Rodriguez, Tamara Carver, Hao Zhang, Tibor Schuster
J Educ Eval Health Prof. 2024;21:31.   Published online November 7, 2024
DOI: https://doi.org/10.3352/jeehp.2024.21.31
  • 4,938 View
  • 257 Download
  • 2 Web of Science
  • 3 Crossref
AbstractAbstract PDFSupplementary Material
Purpose
The primary aim of this study is to validate the Blended Learning Usability Evaluation–Questionnaire (BLUE-Q) for use in the field of health professions education through a Bayesian approach. As Bayesian questionnaire validation remains elusive, a secondary aim of this article is to serve as a simplified tutorial for engaging in such validation practices in health professions education.
Methods
A total of 10 health education-based experts in blended learning were recruited to participate in a 30-minute interviewer-administered survey. On a 5-point Likert scale, experts rated how well they perceived each item of the BLUE-Q to reflect its underlying usability domain (i.e., effectiveness, efficiency, satisfaction, accessibility, organization, and learner experience). Ratings were descriptively analyzed and converted into beta prior distributions. Participants were also given the option to provide qualitative comments for each item.
Results
After reviewing the computed expert prior distributions, 31 quantitative items were identified as having a probability of “low endorsement” and were thus removed from the questionnaire. Additionally, qualitative comments were used to revise the phrasing and order of items to ensure clarity and logical flow. The BLUE-Q’s final version comprises 23 Likert-scale items and 6 open-ended items.
Conclusion
Questionnaire validation can generally be a complex, time-consuming, and costly process, inhibiting many from engaging in proper validation practices. In this study, we demonstrate that a Bayesian questionnaire validation approach can be a simple, resource-efficient, yet rigorous solution to validating a tool for content and item-domain correlation through the elicitation of domain expert endorsement ratings.

Citations

Citations to this article as recorded by  
  • Reliability and construct validation of the Blended Learning Usability Evaluation–Questionnaire with interprofessional clinicians in Canada: a methodological study
    Anish Kumar Arora, Jeff Myers, Tavis Apramian, Kulamakan Kulasegaram, Daryl Bainbridge, Hsien Seow
    Journal of Educational Evaluation for Health Professions.2025; 22: 5.     CrossRef
  • Utilizing cognitive interview in the item refinement of the Blended Teaching Assessment Tool (BTAT) for Health Professions Education
    Maria Teresita B. Dalusong, Glenda Sanggalang Ogerio, Valentin C. Dones, Maria Elizabeth M. Grageda
    Philippine Journal of Health Research and Development.2025; 29(2): 54.     CrossRef
  • All providers Better Communication Skills (ABCs) program: protocol for a randomized controlled trial assessing communication training effectiveness with interprofessional clinicians
    Hsien Seow, Anish K. Arora, Daryl Bainbridge, Zhimeng Jia, Leah Steinberg, Nadia Incardona, Oren Levine, Justin J. Sanders, Jessica Simon, Amanda Roze des Ordons, Karen Zhang, Jeff Myers
    BMC Palliative Care.2025;[Epub]     CrossRef
The effect of simulation-based training on problem-solving skills, critical thinking skills, and self-efficacy among nursing students in Vietnam: a before-and-after study  
Tran Thi Hoang Oanh, Luu Thi Thuy, Ngo Thi Thu Huyen
J Educ Eval Health Prof. 2024;21:24.   Published online September 23, 2024
DOI: https://doi.org/10.3352/jeehp.2024.21.24
  • 10,327 View
  • 575 Download
  • 10 Web of Science
  • 16 Crossref
AbstractAbstract PDFSupplementary Material
Purpose
This study investigated the effect of simulation-based training on nursing students’ problem-solving skills, critical thinking skills, and self-efficacy.
Methods
A single-group pretest and posttest study was conducted among 173 second-year nursing students at a public university in Vietnam from May 2021 to July 2022. Each student participated in the adult nursing preclinical practice course, which utilized a moderate-fidelity simulation teaching approach. Instruments including the Personal Problem-Solving Inventory Scale, Critical Thinking Skills Questionnaire, and General Self-Efficacy Questionnaire were employed to measure participants’ problem-solving skills, critical thinking skills, and self-efficacy. Data were analyzed using descriptive statistics and the paired-sample t-test with the significance level set at P<0.05.
Results
The mean score of the Personal Problem-Solving Inventory posttest (127.24±12.11) was lower than the pretest score (131.42±16.95), suggesting an improvement in the problem-solving skills of the participants (t172=2.55, P=0.011). There was no statistically significant difference in critical thinking skills between the pretest and posttest (P=0.854). Self-efficacy among nursing students showed a substantial increase from the pretest (27.91±5.26) to the posttest (28.71±3.81), with t172=-2.26 and P=0.025.
Conclusion
The results suggest that simulation-based training can improve problem-solving skills and increase self-efficacy among nursing students. Therefore, the integration of simulation-based training in nursing education is recommended.

Citations

Citations to this article as recorded by  
  • Amsterdam Self‐Efficacy Scale for Tooth Removal (ASES‐TR)
    Maaike G. Beuling, Jason Nak, Jens Kober, Jean Pierre T. F. Ho, Jan de Lange, Raoul P. P. P. Grasman, Tom C. T. van Riet
    European Journal of Dental Education.2026; 30(2): 609.     CrossRef
  • Future-ready nurses: Embedding leadership in undergraduate curricula
    Joshua M. Ogle, Nancy A. Claus, Laura Steadman
    Teaching and Learning in Nursing.2026; 21(1): e332.     CrossRef
  • From education to practice: How attitudes shape communication skills in nursing students
    Paola Del Sette, Gabriella Biffa, Annamaria Bagnasco, Gianluca Catania, Milko Zanini, Francesca Riccardi
    Teaching and Learning in Nursing.2026; 21(2): e689.     CrossRef
  • Artificial intelligence self-efficacy and attitudes among nursing students: a multicenter network analysis of educational stratification
    Qin Zeng, Jun Zhu, Yuji Wang, Shaoyu Su, Yan Huang
    BMC Medical Education.2026;[Epub]     CrossRef
  • Career commitment and sense of calling in nursing students: a longitudinal study
    Tingting Ding, Jiaqi Shi, Wenjing Li, Feifei Xiong, Huihui Xu, Chengjia Zhao, Guohua Zhang, Chen Cong
    Frontiers in Medicine.2026;[Epub]     CrossRef
  • Nurse Educators’ Perspectives on Nursing Students’ Critical Thinking Skills and Barriers for Effective Teaching: A Descriptive Cross‐Sectional Study
    Shereen Ragab Dorgham, Friyal Mubark Alqahtani, A. Sana Al-Mahmoud, A. Eshtiaq Al-Faraj, Jordan Tovera Salvador, Lilibeth Dela Victoria Reyes, Kathlynn Buenaobra Sanchez, Basim Mohammed Alanazi, Ahrjaynes Balanag Rosario, Sulaiman Al Sabei
    Nursing Forum.2026;[Epub]     CrossRef
  • RSSDI IJDDC position statement on embracing artificial intelligence responsibly in medical research and scientific publication: consensus recommendations for authors, editors, and peer reviewers
    Rajeev Chawla, Anuj Maheshwari, Alok Modi, Harsh Atul Hirani, Shambo Samrat, Amit Dey, Rakesh Parikh, Banshi Saboo, Manoj Chawla, Amit Gupta, Shalini Jaggi, Sunil Gupta, Sanjay Agarwal, Purvi Chawla, Bharat Saboo, Jothydev Kesavadev, Krishna Seshadri, Awa
    International Journal of Diabetes in Developing Countries.2026; 46(2): 363.     CrossRef
  • Clinical practice experiences of nursing students in Türkiye: a qualitative study using reflexive thematic analysis
    Mehmet Emin Atay, Ramazan Deniz, Bahar Çiftçi
    BMJ Open.2026; 16(5): e118612.     CrossRef
  • Undergraduate nursing students’ learning experiences with mixed reality-based simulation using a complex older adult stroke scenario: A qualitative descriptive study
    Jiyeon Lee
    The Journal of Korean Academic Society of Nursing Education.2026; 32(2): 188.     CrossRef
  • The effects of reciprocal peer practice with digital platform on undergraduate nursing students’ operational skills, critical thinking, and self-efficacy: a quasi-experimental study
    Linfang Xu, Lan Zeng, Chunjuan Liu, Jiayi Zhu, Mengying Qiu, Fengying Zhang
    BMC Medical Education.2026;[Epub]     CrossRef
  • The Effect of Work-Based Learning on Employability Skills: The Role of Self-Efficacy and Vocational Identity
    Suyitno Suyitno, Muhammad Nurtanto, Dwi Jatmoko, Yuli Widiyono, Riawan Yudi Purwoko, Fuad Abdillah, Setuju Setuju, Yudan Hermawan
    European Journal of Educational Research.2025; 14(1): 309.     CrossRef
  • Interactive Success: Empowering Young Minds through Games-Based Learning at NADI PPR Intan Baiduri
    Mohamad Zaki Mohamad Saad, Shafinah Kamarudin, Zuraini Zukiffly, Siti Soleha Zuaimi
    Progress in Computers and Learning .2025; 2(1): 29.     CrossRef
  • Realidad virtual y simulación clínica en la formación de enfermería: impacto en la educación y el desarrollo de habilidades clínicas
    Mario Roberto Sate, María Eugenia Gonzalez, Carmen Graciela Mezacapo, Pablo Andrés Salgado, Gloria Ester Rodríguez
    LATAM Revista Latinoamericana de Ciencias Sociales y Humanidades.2025;[Epub]     CrossRef
  • Utilization of Pusdic Board Game in Increasing Level of Mastery of Grade 8 Learners in Predicting Phenotypic Expressions of Traits
    Charlie T. Anselmo, Jeniffer L. Opeña, Precila C. Delima, Rhea P. Balbi
    International Journal of Multidisciplinary: Applied Business and Education Research.2025; 6(8): 2892.     CrossRef
  • Abusive supervision and nurses’ deviant behaviors: the moderating effect of self-efficacy
    Senay Yürür, Oktay Koç, H. Rıdvan Yurtseven
    BMC Nursing.2025;[Epub]     CrossRef
  • The relationship between nurses’ counseling and problem-solving skills and their job satisfaction
    Mehmet Hayrullah Öztürk, Serap Parlar Kılıç
    Adıyaman Üniversitesi Sağlık Bilimleri Dergisi.2025; 11(3): 283.     CrossRef
Review
Immersive simulation in nursing and midwifery education: a systematic review  
Lahoucine Ben Yahya, Aziz Naciri, Mohamed Radid, Ghizlane Chemsi
J Educ Eval Health Prof. 2024;21:19.   Published online August 8, 2024
DOI: https://doi.org/10.3352/jeehp.2024.21.19
  • 22,052 View
  • 877 Download
  • 16 Web of Science
  • 25 Crossref
AbstractAbstract PDFSupplementary Material
Purpose
Immersive simulation is an innovative training approach in health education that enhances student learning. This study examined its impact on engagement, motivation, and academic performance in nursing and midwifery students.
Methods
A comprehensive systematic search was meticulously conducted in 4 reputable databases—Scopus, PubMed, Web of Science, and Science Direct—following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses guidelines. The research protocol was pre-registered in the PROSPERO registry, ensuring transparency and rigor. The quality of the included studies was assessed using the Medical Education Research Study Quality Instrument.
Results
Out of 90 identified studies, 11 were included in the present review, involving 1,090 participants. Four out of 5 studies observed high post-test engagement scores in the intervention groups. Additionally, 5 out of 6 studies that evaluated motivation found higher post-test motivational scores in the intervention groups than in control groups using traditional approaches. Furthermore, among the 8 out of 11 studies that evaluated academic performance during immersive simulation training, 5 reported significant differences (P<0.001) in favor of the students in the intervention groups.
Conclusion
Immersive simulation, as demonstrated by this study, has a significant potential to enhance student engagement, motivation, and academic performance, surpassing traditional teaching methods. This potential underscores the urgent need for future research in various contexts to better integrate this innovative educational approach into nursing and midwifery education curricula, inspiring hope for improved teaching methods.

Citations

Citations to this article as recorded by  
  • Teaching methods and creative motivation among nursing students: Implications for pedagogical competencies in nursing education
    Gorica Popovska Nalevska, Gordana Ristevska Dimitrovska
    International Journal of Research Studies in Education.2026;[Epub]     CrossRef
  • A Review of Reviews on Virtual Reality in Educational Context: Constraints, Implications, and Research Agendas
    Yuchun Zhong, Chunqi Li, Juming Jiang, Luke Kutszik Fryer, Alex Shum
    Journal of Educational Computing Research.2026; 64(2): 439.     CrossRef
  • Digital Interventions for Palliative Care Education for Nursing Students: A Systematic Review
    Abdulelah Alanazi, Gary Mitchell, Fadwa Naji Al Halaiqa, Fadi Khraim, Stephanie Craig
    Nursing Reports.2026; 16(1): 16.     CrossRef
  • Artificial intelligence self-efficacy and attitudes among nursing students: a multicenter network analysis of educational stratification
    Qin Zeng, Jun Zhu, Yuji Wang, Shaoyu Su, Yan Huang
    BMC Medical Education.2026;[Epub]     CrossRef
  • Escape room and high-fidelity clinical simulation in neonatal nursing education: A randomized clinical trial
    Esperanza Santano-Mogena, Cristina Franco-Antonio, Sergio Cordovilla-Guardia
    Clinical Simulation in Nursing.2026; 114: 101938.     CrossRef
  • The effect of scenario-based high-fidelity simulation on midwifery students’ knowledge and self-efficacy in electronic fetal monitoring: A randomized controlled trial
    İffet Güler Kaya, Neriman Zengin
    Clinical Simulation in Nursing.2026; 113: 101923.     CrossRef
  • Can Virtual Reality Promote Cognitive, Affective, and Psychomotor Learning? A Synthesis of Systematic Reviews and Meta-Analytic Evidence
    Yuchun Zhong, Chunqi Li, Juming Jiang, Luke K. Fryer, Alex Shum
    Journal of Educational Computing Research.2026; 64(6): 1630.     CrossRef
  • Improving Midwifery Students’ Communication and Counseling Skills in Maternal Nutrition: A Quality Improvement Educational Initiative
    Martina Caglioni, Vanessa Odelli, Paolo Ghisleni, Viola Marostica, Elena Celi, Sara Trapani, Barbara Cigoli, Stefania Rinaldi, Stefano Salvatore, Massimo Candiani
    Journal of Midwifery & Women's Health.2026;[Epub]     CrossRef
  • Stepping into the emergency: nursing students’ lived experience in cardiac arrest immersive simulation
    Anna Arnone, Nunzia Pannone, Alice D’Abramo, Francesco Riccardo, Luisa Zeppetella, Alessandro Marino, Giovanni Gioiello
    Teaching and Learning in Nursing.2026; 21(3): e1164.     CrossRef
  • Ebelikte Teknoloji Kullanımı ile İlgili Lisansüstü Tezlerin İncelenmesi: Sistematik Derleme
    Meserret Aslan, Şehma Şen
    Fenerbahçe Üniversitesi Sağlık Bilimleri Dergisi.2026;[Epub]     CrossRef
  • Simulation-Enhanced Learning in Nuclear Medicine: Theory, Modalities, and Applications Across the Training Continuum
    Geoffrey M. Currie, Tarni Nelson, Kym Barry, Johnathan Hewis, David Gilmore
    Journal of Nuclear Medicine Technology.2026; 54(3): 260.     CrossRef
  • Bridging the Gap: The Transition Toward University-Level Midwifery Education and Autonomous Practice in Ukraine
    Igor Lakhno, Victoria Romaieva, Svitlana Pak, Iryna Sykal
    Cureus.2026;[Epub]     CrossRef
  • Comparing the effects of GenAI-assisted and facilitator-led debriefing on knowledge and non-technical skills in immersive virtual reality simulation in nursing education: A quasi-experimental study
    Patrick Pui Kin Kor, Kitty Chan, Alex Pak Lik Tsang, Justina Yat Wa Liu, Kin Cheung, Kam Hung Lai, Elaine Shuk Ping Cheung, Lexi Han Zhi Tan, Hao Chong He
    Nurse Education Today.2026; 166: 107275.     CrossRef
  • The ‘error room’: an innovative approach to safety education for midwives in Morocco
    Latifa Mochhoury, Khaddouj El goundali, Lalla Asmaa Katir Masnaoui, Nabila Msatfa, Milouda Chebabe
    British Journal of Midwifery.2026; 34(8): 444.     CrossRef
  • The Mediating Role of Disgust Sensitivity in the Relationship Between Professional Perception and Professional Belonging Among Midwifery Students: A Cross‐Sectional Study
    Betül Kaplan, Songül Kekil, Ayşe Elkoca
    Nursing Open.2026;[Epub]     CrossRef
  • The power of moulage: Teaching with wound moulage simulations in nursing education
    Romaine Meichtry, Rabea Krings, Alina Klein, Monika Droz, Claudia Schlegel
    Teaching and Learning in Nursing.2025; 20(3): e745.     CrossRef
  • NursingXR: Advancing Nursing Education Through Virtual Reality-Based Training
    Mohammad F. Obeid, Ahmed Ewais, Mohammad R. Asia
    Applied Sciences.2025; 15(6): 2949.     CrossRef
  • Simulation-Based Education and Nursing Student Learning Motivations: A Scoping Review
    Keisuke Nojima, Ryosuke Yoshida, Seiichi Sato, Tomohide Fukuda, Kanae Kakinuma
    Cureus.2025;[Epub]     CrossRef
  • Realidad virtual y simulación clínica en la formación de enfermería: impacto en la educación y el desarrollo de habilidades clínicas
    Mario Roberto Sate, María Eugenia Gonzalez, Carmen Graciela Mezacapo, Pablo Andrés Salgado, Gloria Ester Rodríguez
    LATAM Revista Latinoamericana de Ciencias Sociales y Humanidades.2025;[Epub]     CrossRef
  • The effect of immersive simulation-based learning on an anatomy program in nursing education: a quasi-experimental study
    Lahoucine Ben Yahya, Mohamed Radid, Mohamed El Yaagoubi, Lahcen Elmoumou, Otmane Abouri, Aziz Naciri, Ghizlane Chemsi
    Korean Journal of Medical Education.2025; 37(3): 281.     CrossRef
  • Mixed reality versus manikins in basic life support simulation-based training for medical students in France: the mixed reality non-inferiority randomized controlled trial
    Sofia Barlocco De La Vega, Evelyne Guerif-Dubreucq, Jebrane Bouaoud, Myriam Awad, Léonard Mathon, Agathe Beauvais, Thomas Olivier, Pierre-Clément Thiébaud, Anne-Laure Philippon
    Journal of Educational Evaluation for Health Professions.2025; 22: 15.     CrossRef
  • Examining jigsaw, metaverse, and classic learning methods' impact on nursing students' self-efficacy in pediatric drug administration: a mixed-methods study
    Şerife Tutar, Hande Özgörü, Faruk Durna, Yurdagül Şahin
    Journal of Health Sciences and Medicine.2025; 8(6): 1065.     CrossRef
  • Simulation-Based Learning in Nursing Education: Current Evidence and Future Directions
    Firas Khraisat, Ghada M. Khrais
    Inquisiva Open.2025;[Epub]     CrossRef
  • AI-guided interview simulation to improve employability and reduce anxiety in final-year nursing and midwifery students: A quasi-experimental study
    Seda Sarıköse, Tuba Sengul, Betül Uncu, Holly Kirkland-Kyhn, Nurten Kaya
    Clinical Simulation in Nursing.2025; 109: 101852.     CrossRef
  • Application of Virtual Reality, Artificial Intelligence, and Other Innovative Technologies in Healthcare Education (Nursing and Midwifery Specialties): Challenges and Strategies
    Galya Georgieva-Tsaneva, Ivanichka Serbezova, Silvia Beloeva
    Education Sciences.2024; 15(1): 11.     CrossRef
Research article
Performance of GPT-3.5 and GPT-4 on standardized urology knowledge assessment items in the United States: a descriptive study  
Max Samuel Yudovich, Elizaveta Makarova, Christian Michael Hague, Jay Dilip Raman
J Educ Eval Health Prof. 2024;21:17.   Published online July 8, 2024
DOI: https://doi.org/10.3352/jeehp.2024.21.17
  • 9,744 View
  • 373 Download
  • 22 Web of Science
  • 23 Crossref
AbstractAbstract PDFSupplementary Material
Purpose
This study aimed to evaluate the performance of Chat Generative Pre-Trained Transformer (ChatGPT) with respect to standardized urology multiple-choice items in the United States.
Methods
In total, 700 multiple-choice urology board exam-style items were submitted to GPT-3.5 and GPT-4, and responses were recorded. Items were categorized based on topic and question complexity (recall, interpretation, and problem-solving). The accuracy of GPT-3.5 and GPT-4 was compared across item types in February 2024.
Results
GPT-4 answered 44.4% of items correctly compared to 30.9% for GPT-3.5 (P<0.00001). GPT-4 (vs. GPT-3.5) had higher accuracy with urologic oncology (43.8% vs. 33.9%, P=0.03), sexual medicine (44.3% vs. 27.8%, P=0.046), and pediatric urology (47.1% vs. 27.1%, P=0.012) items. Endourology (38.0% vs. 25.7%, P=0.15), reconstruction and trauma (29.0% vs. 21.0%, P=0.41), and neurourology (49.0% vs. 33.3%, P=0.11) items did not show significant differences in performance across versions. GPT-4 also outperformed GPT-3.5 with respect to recall (45.9% vs. 27.4%, P<0.00001), interpretation (45.6% vs. 31.5%, P=0.0005), and problem-solving (41.8% vs. 34.5%, P=0.56) type items. This difference was not significant for the higher-complexity items.
Conclusions
ChatGPT performs relatively poorly on standardized multiple-choice urology board exam-style items, with GPT-4 outperforming GPT-3.5. The accuracy was below the proposed minimum passing standards for the American Board of Urology’s Continuing Urologic Certification knowledge reinforcement activity (60%). As artificial intelligence progresses in complexity, ChatGPT may become more capable and accurate with respect to board examination items. For now, its responses should be scrutinized.

Citations

Citations to this article as recorded by  
  • Response to letter to the editor Re: Advancements in large language model accuracy for answering physical medicine and rehabilitation board review questions
    Jason Bitterman, Alexander D'Angelo, Alexandra Holachek, James E. Eubanks
    PM&R.2026; 18(1): 109.     CrossRef
  • Assessing the utility of a natural language processing model in answering common urological questions
    Wyatt MacNevin, John‐David Brown, Nicholas Dawe, Jesse T. R. Spooner, Nicholas R. Paterson, Daniel T. Keefe, David G. Bell
    UroPrecision.2026; 4(2): 98.     CrossRef
  • Artificial Intelligence as a Drug Information Resource: Limitations and Strategies to Optimize in Pharmacy Practice
    Christopher Soujah, Carole Bejjani, Nour Adra, Laura Blackburn
    Hospital Pharmacy.2026; 61(2): 117.     CrossRef
  • Assessing multiple chatGPT versions on novel content in the Taiwan urology board examination: accuracy, speed, and domain-specific performance
    Ho Li, Hui-Kung Ting, Yu-Cing Jhuo, Chin-Li Chen, Chien-Chang Kao, Ming-Hsin Yang, Chih-Wei Tsao, En Meng, Sheng-Tang Wu, Pei-Jhang Chiang
    World Journal of Urology.2026;[Epub]     CrossRef
  • Large Language Models Provide Accurate but Potentially Unsafe Answers to Multimodal Critical Care Medicine Board Review Questions
    Ish Sethi, Sharaf Khan, Patrick G. Lyons, Catherine A. Gao, Yuan Luo, Danielle Miltz, Santiago Tovar, Caitlin ten Lohuis, Derek M. Polly, Sagar B. Dave, Michael Sterling, Juan C. Rojas, Xuan Han, Ankit Sakhuja, Leo A. Celi, Greg S. Martin, Craig M. Cooper
    Critical Care Medicine.2026;[Epub]     CrossRef
  • Evaluating the Performance of ChatGPT4.0 Versus ChatGPT3.5 on the Hand Surgery Self-Assessment Exam: A Comparative Analysis of Performance on Image-Based Questions
    Kiera L Vrindten, Megan Hsu, Yuri Han, Brian Rust, Heili Truumees, Brian M Katt
    Cureus.2025;[Epub]     CrossRef
  • Assessing the performance of large language models (GPT-3.5 and GPT-4) and accurate clinical information for pediatric nephrology
    Nadide Melike Sav
    Pediatric Nephrology.2025; 40(9): 2879.     CrossRef
  • Retrieval-augmented generation enhances large language model performance on the Japanese orthopedic board examination
    Juntaro Maruyama, Satoshi Maki, Takeo Furuya, Yuki Nagashima, Kyota Kitagawa, Yasunori Toki, Shuhei Iwata, Megumi Yazaki, Takaki Kitamura, Sho Gushiken, Yuji Noguchi, Masataka Miura, Masahiro Inoue, Yasuhiro Shiga, Kazuhide Inage, Sumihisa Orita, Seiji Oh
    Journal of Orthopaedic Science.2025; 30(6): 1193.     CrossRef
  • Advancements in large language model accuracy for answering physical medicine and rehabilitation board review questions
    Jason Bitterman, Alexander D'Angelo, Alexandra Holachek, James E. Eubanks
    PM&R.2025; 17(9): 1091.     CrossRef
  • Accuracy of Large Language Models When Answering Clinical Research Questions: Systematic Review and Network Meta-Analysis
    Ling Wang, Jinglin Li, Boyang Zhuang, Shasha Huang, Meilin Fang, Cunze Wang, Wen Li, Mohan Zhang, Shurong Gong
    Journal of Medical Internet Research.2025; 27: e64486.     CrossRef
  • OpenAI o1 Large Language Model Outperforms GPT-4o, Gemini 1.5 Flash, and Human Test Takers on Ophthalmology Board–Style Questions
    Ryan Shean, Tathya Shah, Sina Sobhani, Alan Tang, Ali Setayesh, Kyle Bolo, Van Nguyen, Benjamin Xu
    Ophthalmology Science.2025; 5(6): 100844.     CrossRef
  • A comparative analysis of DeepSeek R1, DeepSeek-R1-Lite, OpenAi o1 Pro, and Grok 3 performance on ophthalmology board-style questions
    Ryan Shean, Tathya Shah, Aditya Pandiarajan, Alan Tang, Kyle Bolo, Van Nguyen, Benjamin Xu
    Scientific Reports.2025;[Epub]     CrossRef
  • The performance of ChatGPT on medical image-based assessments and implications for medical education
    Xiang Yang, Wei Chen
    BMC Medical Education.2025;[Epub]     CrossRef
  • Applications, Challenges, and Prospects of Generative Artificial Intelligence Empowering Medical Education: Scoping Review
    Yuhang Lin, Zhiheng Luo, Zicheng Ye, Nuoxi Zhong, Lijian Zhao, Long Zhang, Xiaolan Li, Zetao Chen, Yijia Chen
    JMIR Medical Education.2025; 11: e71125.     CrossRef
  • ChatGPT’s role in the rapidly evolving hematologic cancer landscape
    Tiffany Nong, Sean Britton, Viralkumar Bhanderi, Justin Taylor
    Future Science OA.2025;[Epub]     CrossRef
  • Performance of ChatGPT-4 on the French Board of Plastic Reconstructive and Aesthetic Surgery written exam: a descriptive study
    Emma Dejean-Bouyer, Anoujat Kanlagna, François Thuau, Pierre Perrot, Ugo Lancien
    Journal of Educational Evaluation for Health Professions.2025; 22: 27.     CrossRef
  • Technologies, opportunities, challenges, and future directions for integrating generative artificial intelligence into medical education: a narrative review
    Junseok Kang, Jihyun Ahn
    Ewha Medical Journal.2025; 48(4): e53.     CrossRef
  • Performance of the ChatGPT-5 Language Model in Solving a Specialty Examination in Balneology and Physical Medicine
    Michalina Loson-Kawalec, Anna Kowalczyk, Dawid Szymanski, Patrycja Dadynska, Aleksander Tabor, Dawid Bartosik , Marta Zerek, Gracjan Sitarek, Bartosz Starzynski, Alina Keska, Bartlomiej Cwikla, Piotr Sawina, Tomasz Dolata, Adrianna Pielech, Maciej Majchrz
    Cureus.2025;[Epub]     CrossRef
  • Diagnostic accuracy and bias in open access and subscription-based large language models for multiple sclerosis and neuromyelitis optica spectrum disorder
    Tom G. Punnen, Kevin S. Shan, Mahi A. Patel, Morgan C. McCreary, Diem H. Tran, Jose R. Santoyo, Katy W. Burgess, Tatum M. Moog, Alexander D. Smith, Darin T. Okuda
    Intelligence-Based Medicine.2025; 12: 100314.     CrossRef
  • Privacy-by-Design Framework for Large Language Model Chatbots in Urology
    Eun Joung Kim, JungYoon Kim
    International Neurourology Journal.2025; 29(Suppl 2): S65.     CrossRef
  • Potential and pitfalls: accuracy versus adequacy of ChatGPT’s performance on surgery shelf examination
    Baylee Brochu, Michael D. Cobler-Lichter, Talia R. Arcieri, Nikita M. Shah, Jessica M. Delamater, Ana M. Reyes, Matthew S. Sussman, Edward B. Lineen, Laurence R. Sands, Vanessa W. Hui, Steven E. Rodgers, Chad M. Thorson
    Global Surgical Education - Journal of the Association for Surgical Education.2025;[Epub]     CrossRef
  • From GPT-3.5 to GPT-4.o: A Leap in AI’s Medical Exam Performance
    Markus Kipp
    Information.2024; 15(9): 543.     CrossRef
  • Artificial Intelligence can Facilitate Application of Risk Stratification Algorithms to Bladder Cancer Patient Case Scenarios
    Max S Yudovich, Ahmad N Alzubaidi, Jay D Raman
    Clinical Medicine Insights: Oncology.2024;[Epub]     CrossRef

JEEHP : Journal of Educational Evaluation for Health Professions
TOP