Abstract
The Diagnostic Competency During Simulation-Based (DCDS) Learning Tool (Burt & Olson, 2023) is an advanced pedagogical and psychometric instrument engineered to systematically evaluate individual diagnostic reasoning competencies among nurse practitioner (NP) students participating in simulation-based learning experiences. Developed in response to national imperatives regarding patient safety and the reduction of diagnostic error, the DCDS Learning Tool operationalizes diagnostic cognition into observable, behavioral indicators. The instrument comprises 28 items organized across six discrete competency domains derived from established clinical frameworks and curriculum objectives (Ahmed et al., 2017; American Board of Internal Medicine [ABIM], 2021). Evaluation of the tool incorporates an ordinal three-tier scoring paradigm (minimal, partial, and complete) that requires raters to gauge scenario-specific contextual factors before assessing observed performance. Psychometric evaluation of the instrument demonstrates robust content validity, with individual competency domain Scale Content Validity Index (S-CVI) scores ranging from 0.9175 to 1.00 and an overall total S-CVI of 0.98. Assessment of interrater reliability yielded an intraclass correlation coefficient (ICC) of 0.548 (95% confidence interval [CI: 0.482, 0.612], p < 0.0001), indicating fair-to-moderate consistency among faculty raters under initial testing conditions. The DCDS Learning Tool serves both formative and summative educational purposes, offering clinical nurse educators granular, actionable diagnostic metrics that inform targeted debriefing, remediate cognitive vulnerabilities, and foster diagnostic expertise in healthcare trainees.
Keywords
Diagnostic reasoning, nurse practitioner education, simulation-based learning, clinical competency, diagnostic error, assessment tool, educational measurement, dual-process theory, psychometrics, content validity index, interrater reliability, formative assessment, clinical judgment
Authors
The Diagnostic Competency During Simulation-Based Learning Tool was conceptualized, designed, and psychometrically validated by:
- Leah Burt, PhD, APRN, ANP-BC — Department of Biobehavioral Nursing Science, College of Nursing, University of Illinois Chicago, Chicago, Illinois, United States. Correspondence: [email protected].
- Andrew Olson, MD, FACP, FAAP — Departments of Medicine and Pediatrics, University of Minnesota Medical School, Minneapolis, Minnesota, United States.
Purpose
Diagnostic errors represent one of the most pervasive, harmful, and costly threats to patient safety across global healthcare systems. According to landmark investigations by the National Academies of Sciences, Engineering, and Medicine (NASEM, 2015), most individuals will experience at least one diagnostic error in their lifetime, sometimes with catastrophic consequences. Despite these realities, formal training and rigorous, standardized assessment of diagnostic reasoning competencies have historically been marginalized or inconsistently addressed within advanced practice nursing curricula. Nurse practitioner students frequently acquire clinical problem-solving skills via implicit, unstandardized experiential learning, leaving cognitive biases, premature closure, and suboptimal diagnostic strategies unaddressed.
The overarching purpose of the DCDS Learning Tool is to transform diagnostic reasoning assessment from a subjective, holistic impression into an objective, standardized, and competency-aligned measurement process. Specifically, the instrument seeks to:
- Provide nurse practitioner educators with a psychometrically sound, granular mechanism to evaluate observable behaviors that reflect internal diagnostic decision-making during high-fidelity, standardized patient, or virtual simulation-based clinical scenarios.
- Facilitate actionable formative feedback during simulation debriefings by identifying specific vulnerabilities within the diagnostic trajectory, such as failing to generate a comprehensive differential diagnosis, neglecting to prioritize life-threatening conditions, or succumbing to cognitive biases like confirmation bias or anchoring.
- Establish a standardized framework for summative clinical evaluations, ensuring that graduating nurse practitioners possess the minimum threshold of diagnostic competence necessary for safe, independent prescriptive and clinical practice.
- Equip educational researchers with an empirical instrument to quantify the efficacy of diagnostic reasoning interventions, clinical reasoning curricula, and novel simulation modalities across diverse academic and clinical settings.
Psychological Construct
The primary psychological construct captured by the DCDS Learning Tool is individual diagnostic reasoning competency. Within cognitive psychology and healthcare education, diagnostic reasoning is defined as the complex, iterative, and dynamic cognitive process through which a clinician continuously acquires, organizes, synthesizes, and interprets clinical data to formulate an accurate diagnosis and subsequent management plan. Rather than functioning as a unitary trait, diagnostic competency encompasses a constellation of metacognitive abilities, domain-specific clinical knowledge structures, information-seeking behaviors, and socio-communicative skills.
The DCDS Learning Tool operationalizes diagnostic reasoning through 28 distinct behavioral items distributed across six interconnected competency domains derived from the literature on diagnostic excellence (Ahmed et al., 2017; ABIM, 2021):
- Domain 1: Information Gathering and Clinical Exploration — Evaluates the learner’s ability to execute a hypothesis-directed history and physical examination. Competent performance involves asking precise, high-yield questions, performing appropriate physical assessment maneuvers, and actively avoiding exploratory ‘shotgun’ questioning that clutters cognitive bandwidth.
- Domain 2: Hypothesis Generation and Differential Formulation — Measures the ability to generate a prioritized, clinically justifiable differential diagnosis early in the clinical encounter. This requires learners to balance high-probability diagnoses (common presentations) against low-probability, high-consequence diagnoses (‘cannot-miss’ conditions).
- Domain 3: Problem Representation and Clinical Synthesis — Assesses the clinician’s capacity to summarize complex, messy clinical data into a succinct, abstract clinical problem representation (frequently termed a ‘one-liner’ or ‘illness script’ summary). This domain captures how effectively the student leverages semantic qualifiers (e.g., acute vs. chronic, unilateral vs. bilateral) to abstract patient findings.
- Domain 4: Targeted Diagnostic Verification and Testing — Measures the deliberate selection, sequencing, and interpretation of diagnostic studies (laboratory, imaging, and bedside testing). It evaluates whether the learner chooses diagnostic tests based on pre-test probability, understands test operating characteristics (sensitivity and specificity), and minimizes low-yield or redundant investigations.
- Domain 5: Cognitive Metacognition, Bias Mitigation, and Reflection — Evaluates the student’s reflective capacity during and after the diagnostic process. Competency is demonstrated when the student pauses to consider alternative explanations, engages in ‘diagnostic time-outs’, explicitly queries whether findings fit the working hypothesis, and actively identifies personal cognitive shortcuts (heuristics) that might distort clinical judgment.
- Domain 6: Diagnostic Communication and Shared Decision-Making — Examines the learner’s ability to communicate diagnostic hypotheses, inherent diagnostic uncertainty, and proposed testing rationale clearly and compassionately to the patient, interprofessional team members, and supervising clinicians.
Theoretical Framework
The DCDS Learning Tool is grounded in the intersection of modern cognitive psychology, expert-performance theory, and simulation pedagogy. Four foundational theoretical models inform the structural and conceptual basis of the instrument:
1. Dual-Process Theory of Cognition
Rooted in the cognitive psychology of Daniel Kahneman (2011) and its medical application by Pat Croskerry (2009), dual-process theory posits that clinical decision-making relies on two distinct modes of thinking: System 1 (fast, intuitive, heuristic-driven, subconscious, and low-effort) and System 2 (slow, analytical, deliberate, conscious, and resource-intensive). Experienced clinicians frequently utilize System 1 pattern matching; however, novice and student nurse practitioners must systematically engage System 2 analytical reasoning to avoid diagnostic error. The DCDS Learning Tool measures observable behaviors that indicate appropriate calibration and switching between System 1 heuristics and System 2 analytical verification, emphasizing metacognitive surveillance to avoid premature closure.
2. Script Theory and Knowledge Encapsulation
As conceptualized by Schmidt and Boshuizen (1993), clinical expertise develops through the formation and refinement of illness scripts—rich, multidimensional mental representations stored in long-term memory that link enabling conditions (predispositions, epidemiology), fault mechanisms (pathophysiology), and clinical consequences (signs, symptoms). The DCDS evaluates how effectively learners access, compare, and contrast competing illness scripts when encountering complex, ambiguous patient presentations in simulation.
3. Situated Cognition and Experiential Learning
Drawing from David Kolb’s (1984) Experiential Learning Theory and situated learning theory (Lave & Wenger, 1991), diagnostic competence cannot be fully assessed via static, paper-based testing because diagnostic thinking is deeply situated within the environmental, social, and emotional context of practice. Simulation provides an authentic, high-fidelity clinical ecology wherein students must manage clinical uncertainty, time pressure, and dynamic patient responses. The DCDS serves as the observational bridge that turns concrete simulation experiences into abstract conceptualizations during debriefing.
4. Cognitive Load Theory
In accordance with John Sweller’s (1988) Cognitive Load Theory, learners possess limited working memory capacity. High-fidelity clinical simulations inherently produce high intrinsic and extraneous cognitive loads. The DCDS assesses how systematically learners manage cognitive load through structured information gathering and deliberate synthesis, allowing educators to identify when cognitive overload triggers breakdown in diagnostic reasoning.
Validity
The initial psychometric evaluation conducted by Burt and Olson (2023) focused heavily on content and construct-related validity evidence, adhering to guidelines established by the Standards for Educational and Psychological Testing (AERA, APA, & NCME, 2014) and Lynn’s (1986) criteria for content validation.
Content Validity
Content validity was evaluated using a panel of national experts in diagnostic reasoning, simulation-based nursing education, and clinical measurement. Experts quantitatively scored the clarity, relevance, and representativeness of each item on a 4-point ordinal scale. Content Validity Index (CVI) metrics were computed at both the individual item level (I-CVI) and scale level (S-CVI):
- Domain-Level S-CVI: Individual competency domain scale content validity index scores demonstrated high alignment with expert judgment, ranging from 0.9175 to 1.00 across all six domains.
- Total Scale S-CVI: The total scale content validity index (S-CVI/Ave) achieved an exceptional score of 0.98, comfortably exceeding the standard benchmark threshold of 0.90 for newly developed educational scales.
- Qualitative Panel Review: Qualitative feedback from expert reviewers guided item refinement, ensuring that terminology precisely reflected clinical reasoning benchmarks established by the Society to Improve Diagnosis in Medicine (SIDM) and the American Board of Internal Medicine Foundation.
Construct and Consequential Validity
Construct validity in the DCDS is supported by its structural derivation from empirically established frameworks (Ahmed et al., 2017). By mapping each of the 28 behavioral items to discrete cognitive benchmarks, the tool demonstrates high construct fidelity. Furthermore, consequential validity—the educational and clinical repercussions of score interpretation—is supported by the tool’s explicit emphasis on actionable feedback. Rather than rendering a punitive global pass/fail verdict, the DCDS generates an actionable behavioral profile that directly guides post-simulation remediation.
Reliability
Evaluating the reliability of observational assessment instruments in clinical simulation presents unique psychometric challenges, as score variance stems not only from student performance but also from rater subjectivity, scenario difficulty, and rater training. Burt and Olson (2023) investigated the interrater reliability (IRR) of the DCDS Learning Tool across trained nursing education faculty members observing simulated student encounters.
- Intraclass Correlation Coefficient (ICC): The overall intraclass correlation coefficient for the DCDS tool was established at 0.548 (p < 0.0001, 95% confidence interval CI [0.482, 0.612]).
- Interpretation of Reliability Metric: According to standard psychometric criteria (Cicchetti, 1994), an ICC value between 0.40 and 0.59 represents fair to moderate interrater agreement. In the context of complex, high-dimensional observational assessments involving multiple raters evaluating fluid clinical behaviors, an ICC of 0.548 represents a statistically significant, viable baseline.
- Sources of Measurement Variance: Psychometric analysis revealed that rater variability was primarily driven by divergent interpretations of the ‘partial’ rating boundary, subtle differences in rater attention to diagnostic communication versus diagnostic testing, and varied clinical backgrounds among faculty raters.
- Recommendations for Maximizing Reliability: To elevate interrater reliability in operational settings, the authors emphasize the necessity of robust Frame-of-Reference (FOR) training for faculty raters prior to tool administration, calibration sessions using pre-recorded benchmark simulations, and clearly articulated scenario-specific rubrics.
Factor Analysis
In the primary validation study by Burt and Olson (2023), no formal exploratory factor analysis (EFA) or confirmatory factor analysis (CFA) data were reported. The initial developmental phase prioritized content validity consensus, behavioral indicator calibration, and preliminary interrater reliability within a focused academic simulation environment.
Conducting robust factor analysis on observational simulation rubrics requires substantial sample sizes—typically a minimum of 5 to 10 respondents per item, which equates to an empirical sample of 140 to 280 standardized student observations for a 28-item scale. In high-fidelity simulation environments, obtaining such sample sizes poses significant resource and scheduling barriers. Consequently, the internal dimensionality of the DCDS tool currently rests upon its theoretical and rational structural derivation from the six published competency domains (Ahmed et al., 2017).
Future psychometric investigations should leverage multi-center simulation cohorts to execute CFA, evaluating whether the 28 items conform to a six-factor oblique model, a single higher-order general diagnostic competence factor, or a bifactor structure that simultaneously accounts for general diagnostic capability alongside domain-specific cognitive proficiencies. Additionally, the application of Many-Facet Rasch Measurement (MFRM) is strongly recommended to simultaneously model learner ability, rater severity, and simulation scenario difficulty.
Instrument / Measurement Tool
The Diagnostic Competency During Simulation-Based (DCDS) Learning Tool is an observational rubric specifically formulated for clinical educators. Below are the structural specifications and operational guidelines of the instrument:
- Instrument Name: Diagnostic Competency During Simulation-Based (DCDS) Learning Tool
- Primary Author Citation: Burt, L., & Olson, A. (2023)
- Construct Assessed: Individual diagnostic reasoning competence in simulation environments
- Target Population: Nurse practitioner (NP) students and advanced practice nursing learners (applicable across adult-gerontology, family, acute care, and pediatric tracks)
- Target Raters: Nursing education faculty, clinical preceptors, simulation educators, and diagnostic reasoning researchers
- Test Format: Standardized observational behavioral rubric designed for real-time or video-recorded simulation evaluation
- Total Item Count: 28 observable behavioral items
- Domain Architecture: Six competency domains (Information Gathering, Hypothesis Generation, Problem Representation, Diagnostic Verification, Metacognition/Bias Mitigation, Diagnostic Communication)
- Response Scale and Rating Rules:
- Authentic Rating Paradigm: Raters assign ordinal-level judgments to observed student behaviors using three standardized performance tiers: minimal, partial, or complete.
- Scenario Contextualization: Raters are explicitly guided to analyze scenario characteristics (e.g., acuity, complexity, diagnostic ambiguity) prior to evaluating performance.
- Conservative Scoring Guideline: If a student demonstrates only a portion of the behaviors described within an item, raters are strictly advised to assign a lower rating (e.g., defaulting to ‘partial’ instead of ‘complete’, or ‘minimal’ instead of ‘partial’).
- Pattern and Discrepancy Analysis: Educators are instructed to review ratings across all domains to identify longitudinal trends, strengths, and inconsistencies between interrelated competencies.
- Feedback Integration: During debriefing sessions, educators focus specifically on the behavioral indicators needed for the learner to advance from their current rating to the next competency tier.
- Language: English
- Time Required: Varies with simulation length (typically 15–30 minutes of direct simulation observation, followed by 5–10 minutes of rubric completion and scoring synthesis)
Permissions & Fee and Test Year
- Test Year: 2023
- Original Copyright: © 2023 Elsevier Inc. on behalf of the American Association of Colleges of Nursing (AACN). Published in the Journal of Professional Nursing.
- Fee: Free of charge for non-commercial educational, academic, and scientific research purposes, subject to proper academic attribution.
- Permissions & Inquiries: Educational institutions, researchers, and simulation directors wishing to implement or adapt the complete 28-item DCDS Learning Tool must seek permission from the corresponding author: Leah Burt, PhD, APRN, ANP-BC, University of Illinois Chicago College of Nursing, Department of Biobehavioral Nursing Science, 845 South Damen Avenue (MC 802), Chicago, IL 60612, USA. Email: [email protected].
References
Ahmed, A., Kothari, A., Bashir, A., & Olson, A. P. (2017). A national assessment of diagnostic reasoning curricula in internal medicine residency programs. Diagnosis, 4(4), 213–218. https://doi.org/10.1515/dx-2017-0019
American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association.
Burt, L., & Olson, A. (2023). Development and psychometric testing of the Diagnostic Competency During Simulation-based (DCDS) Learning Tool. Journal of Professional Nursing, 45, 51–59. https://doi.org/10.1016/j.profnurs.2023.01.008
Cicchetti, D. V. (1994). Guidelines, criteria, and rules of thumb for evaluating normed and standardized assessment instruments in psychology. Psychological Assessment, 6(4), 284–290. https://doi.org/10.1037/1040-3590.6.4.284
Croskerry, P. (2009). A universal model of diagnostic reasoning. Academic Medicine, 84(8), 1022–1028. https://doi.org/10.1097/ACM.0b013e3181ace703
Kahneman, D. (2011). Thinking, fast and slow. Farrar, Straus and Giroux.
Kolb, D. A. (1984). Experiential learning: Experience as the source of learning and development. Prentice-Hall.
Lave, J., & Wenger, E. (1991). Situated learning: Legitimate peripheral participation. Cambridge University Press. https://doi.org/10.1017/CBO9780511815355
Lynn, M. R. (1986). Determination and quantification of content validity. Nursing Research, 35(6), 382–385. https://doi.org/10.1097/00006199-198611000-00017
National Academies of Sciences, Engineering, and Medicine. (2015). Improving diagnosis in health care. The National Academies Press. https://doi.org/10.17226/21794
Schmidt, H. G., & Boshuizen, H. P. (1993). On the origin of intermediate effects in clinical reasoning: The down side of intermediate levels of cognitive structure. Cognitive Science, 17(3), 359–383. https://doi.org/10.1207/s15516709cog1703_3
Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257–285. https://doi.org/10.1207/s15516709cog1202_4