1. Abstract
The Computational Thinking Test for Lower Primary (CTtLP) is an empirically validated psychometric instrument developed by Shuhan Zhang and Gary K. W. Wong (2023) to assess foundational computational thinking (CT) competencies in early childhood and lower elementary school populations, specifically targeting children in Grades 1 through 3 (ages 6 to 10 years). Framed by the rigorous methodological principles of Evidence-Centered Design (ECD), the CTtLP was established to address a pronounced diagnostic gap in early psychoeducational and computer science education: the lack of psychometrically sound, age-appropriate, and standardized objective assessments capable of capturing computational reasoning prior to formal coding instruction. The assessment comprises 27 multiple-choice constructed-response items operationalized across three real-world problem-solving scenarios. Each item presents five distinct response options: one keyed correct response, three carefully engineered distractors designed to capture common developmental cognitive missteps, and an explicit ‘I don’t know’ option introduced to mitigate random guessing and pseudo-guessing bias.
The psychometric evaluation of the instrument was executed through a complementary integration of Classical Test Theory (CTT) and modern Item Response Theory (IRT), using the three-parameter logistic (3PL) model. Rigorous field-testing, cognitive interviews, and expert validation refined an initial pool of 30 items down to the final 27-item instrument. The CTtLP demonstrated high internal consistency reliability, with a Cronbach’s alpha coefficient of 0.873, and solid test-retest reliability across an eight-week stability window with an Intraclass Correlation Coefficient (ICC) of 0.757. Confirmatory factor analysis verified the underlying unidimensionality of the construct (Root Mean Square Error of Approximation [RMSEA] = 0.041; Comparative Fit Index [CFI] = 0.911; Tucker-Lewis Index [TLI] = 0.900). Furthermore, IRT calibrations confirmed excellent item discrimination parameters (mean a = 2.188), balanced difficulty thresholds (mean b = 0.353, ranging from -1.096 to 1.892), and tightly constrained guessing estimates (mean c = 0.136). The instrument exhibits sound criterion-related validity, demonstrating a statistically significant moderate correlation with elementary academic course achievement (r = 0.443, p < 0.001), rendering it an exceptional diagnostic and evaluative tool for psychometricians, educational psychologists, and curriculum designers worldwide.
2. Keywords
Computational Thinking, Lower Primary Students, Evidence-Centered Design, Classical Test Theory, Item Response Theory, Early Childhood Education, Psychoeducational Assessment, Error Identification, Instruction Sequence, Psychometrics
3. Authors
The Computational Thinking Test for Lower Primary (CTtLP) was conceptualized, designed, and psychometrically validated by researchers based at the Faculty of Education within the University of Hong Kong:
- Dr. Shuhan Zhang (ORCID: 0000-0002-6493-9979) — Faculty of Education, The University of Hong Kong, Pokfulam Road, Hong Kong SAR, China. Email: [email protected].
- Dr. Gary K. W. Wong (ORCID: 0000-0003-1269-0734) — Associate Professor, Faculty of Education, The University of Hong Kong, Pokfulam Road, Hong Kong SAR, China.
Correspondence regarding the theoretical derivation, administrative protocol, and licensing permissions for research utilization should be directed to Dr. Shuhan Zhang at the University of Hong Kong.
4. Purpose
The primary purpose of the Computational Thinking Test for Lower Primary (CTtLP) is to serve as an empirically grounded, psychometrically robust, and developmentally aligned diagnostic tool to quantify computational thinking (CT) acquisition among young learners in lower primary education (Grades 1–3, typically aged 6 to 10 years). In recent decades, educational theorists and computer scientists, stimulated by seminal paradigms such as those introduced by Jeannette Wing, have reconceptualized computational thinking not merely as an ancillary technical skill associated with professional computer programming, but as a fundamental, cross-disciplinary literacy on par with reading, writing, and arithmetic. Despite this global curricular shift, psychometric instrumentation historically lagged behind pedagogical intent. Most existing CT assessment batteries were constructed for secondary school pupils, undergraduate computer science majors, or upper primary students (e.g., Grades 4 through 6). These legacy scales regularly presume high reading comprehension abilities, foundational keyboarding fluency, and pre-existing familiarity with formal block-based or syntax-based programming environments such as Scratch, Blockly, or Python.
Consequently, assessing CT capabilities in early childhood has historically faced profound methodological hurdles, including cognitive confounders such as reading load, working memory overload, and varying socio-economic access to digital hardware. The CTtLP directly addresses these limitations by offering a standardized, paper-and-pencil or screen-administered constructed-response multiple-choice inventory that decouples underlying computational logic from syntactical coding proficiency. Rather than asking a seven-year-old child to write scripts or parse technical code, the CTtLP frames computational tasks within relatable, everyday scenarios that evaluate core logical structures: the mental execution of sequential algorithms, the prediction of instruction outputs, and the systematic identification and remediation of algorithmic errors (debugging).
In applied psychoeducational and educational settings, the CTtLP serves critical evaluative functions:
- Curricular Evaluation and Program Benchmarking: It affords educational authorities, school districts, and curriculum architects an objective yardstick to assess the longitudinal efficacy of computer science and STEM interventions introduced in early childhood education.
- Individual Diagnostic Screening: School psychologists and learning specialists can deploy the CTtLP to detect relative cognitive strengths and developmental delays in algorithmic thinking, logical sequencing, and spatial-analytical planning before these deficits manifest in later STEM-related academic underachievement.
- Empirical Research and Psychoeducational Modeling: The instrument facilitates cross-sectional and longitudinal empirical investigations into the neurocognitive substrates of computational thinking, enabling researchers to explore how emerging algorithmic competencies intersect with broader cognitive faculties, including executive functioning, fluid intelligence, working memory, and spatial reasoning.
By establishing an objective metric tailored to the developmental characteristics of 6- to 10-year-olds, the CTtLP ensures that pedagogical diagnostics can take place during the critical developmental window when children transition from preoperational to concrete operational thought.
5. Psychological Construct
The core latent psychological construct measured by the CTtLP is foundational Computational Thinking (CT), defined operationalistically within the sphere of lower primary education as the cognitive capacity to formulate problems, decompose systems, construct logical instruction sequences, simulate the execution of step-by-step algorithms, and systematically detect and rectify errors. Grounded in both psychometric principles and cognitive development theory, the CTtLP operationalizes computational thinking as a unidimensional cognitive ability underlying several interrelated cognitive operational domains:
Algorithmic Sequencing and Instruction Execution
At the lower primary level, algorithmic thinking does not manifest as high-level abstract mathematical coding; rather, it reflects a child’s ability to mentally project, follow, and combine discrete, non-ambiguous operations in a deterministic, chronological sequence. Within the CTtLP, this dimension evaluates how a young learner navigates an ordered series of operations—such as spatial movements (e.g., forward, turn left, turn right), object manipulations, or symbolic transformations. The child must internally coordinate the initial state of a system, track each sequential operation without losing track of state changes, and accurately predict the terminal output. This process demands substantial reliance on the visuospatial sketchpad and central executive components of working memory, requiring the child to preserve intermediate mental representations as the instruction chain unfolds.
Error Identification and Algorithmic Debugging
Algorithmic debugging represents one of the most sophisticated metacognitive dimensions of early computational thinking. It requires the child to engage in mental simulation, discrepancy detection, and deductive reasoning. Rather than executing a sequence forwards to observe an outcome, the child is presented with a designated goal state alongside an imperfect sequence of instructions that fails to achieve that goal. The child must mentally step through the algorithm, identify the exact point of divergence between the intended trajectory and the actual execution, and determine which instruction is redundant, missing, or oriented in the incorrect direction. This capacity requires flexible cognitive shifting, inhibition of impulsive answering, and high levels of error-monitoring competence.
Decomposition and Pattern Abstraction within Concrete Scenarios
To render these advanced computational processes accessible to lower primary students without incurring excessive cognitive load, the CTtLP embeds tasks into three accessible, real-world narrative scenarios. These scenarios demand concrete problem decomposition: breaking down a broader navigational or manipulative goal into isolated sub-problems. Children must abstract the salient directional and logical patterns while ignoring extraneous visual noise. Distractors across the test were engineered based on common developmental error patterns, such as reversing egocentric versus intrinsic spatial references (e.g., confusing the character’s ‘left’ with the child’s ‘left’) or failing to anticipate cumulative step-by-step compound rotations.
6. Theoretical Framework
The architectural foundation of the Computational Thinking Test for Lower Primary is rooted in two intersecting theoretical frameworks: the assessment engineering framework of Evidence-Centered Design (ECD) and the developmental cognitive paradigms of Jean Piaget and Seymour Papert.
Evidence-Centered Design (ECD)
Developed by Mislevy, Steinberg, and Almond (2003), Evidence-Centered Design provides a formal, evidentiary argument structure that guides test developers to ensure valid inferences about a test-taker’s unobservable psychological competencies based on observable performance. The CTtLP operationalizes the foundational layers of the ECD framework:
- The Competency Model (Student Model): Defines the unobservable latent trait ($ heta$)—specifically, computational thinking ability in early elementary childhood, characterized by instruction interpretation, sequential execution, and algorithmic debugging.
- The Evidence Model: Delineates the behaviors and response choices that provide evidential warrants for the target competencies. It identifies what constitutes observable proof of algorithmic understanding versus specific misconceptions. In the CTtLP, selecting the single keyed answer across diverse item contexts serves as positive evidence, while the systematic selection of specific distractors provides empirical evidence of distinct computational errors (e.g., rotational orientation errors, premature termination of steps).
- The Task Model: Specifies the structural conditions and visual scenario parameters under which tasks are administered. The CTtLP designed three standard, context-rich scenarios that maintain semantic simplicity, visual clarity, and minimal linguistic demands, thus preventing reading comprehension from becoming a confounding variance source.
Developmental Cognitive Foundations: Piaget and Papert’s Constructionism
The design of the CTtLP aligns precisely with Piaget’s stage theory of cognitive development. Children in Grades 1 through 3 (ages 6 to 10) are undergoing the critical developmental transition from the preoperational stage to the concrete operational stage. During this period, children gradually develop operational schemas: they acquire the ability to conserve quantity, understand spatial relations, internalize mental reversibility, and coordinate external reference frames. However, their cognitive operations remain tethered to concrete, tangible representations rather than purely formal, abstract propositional logic.
Concurrently, the CTtLP draws upon the Constructionist learning theory pioneered by Seymour Papert. Papert demonstrated through the LOGO Turtle programming environment that young children can master powerful computational concepts if those concepts are externalized through manipulable, “object-to-think-with” frameworks. The CTtLP embeds these principles by presenting tasks where algorithmic paths and instructions are visually externalized as spatial movements, arrow sequences, and tangible transformations, thereby allowing children to apply concrete operational logic to formal computational problems.
7. Validity
The validation process of the CTtLP followed a multi-phased psychometric protocol designed to establish content, construct, and criterion-related validity evidence in alignment with the Standards for Educational and Psychological Testing (AERA, APA, & NCME).
Content Validity and Item Evolution
Initial content validation was executed through a rigorous four-stage iterative design process:
- Expert Panel Review: An expert committee comprising educational psychologists, early childhood education specialists, and computer science faculty evaluated an initial pool of 30 items for conceptual alignment with ECD competency frameworks, developmental appropriateness, visual clarity, and linguistic simplicity.
- Cognitive Interviews: “Think-aloud” protocol interviews were conducted with individual lower primary children to directly observe how young learners parsed the visual prompts, interpreted task demands, and reasoned through the choices. This phase uncovered minor ambiguities in graphical icon orientation, leading to refinements in item artwork.
- Pilot and Field Testing: An empirical field test evaluated preliminary item characteristics, identifying problematic items that exhibited severe ceiling/floor effects, negative point-biserial correlations, or visual misinterpretations.
- Final Item Selection: Following these iterations, 27 psychometrically optimal items were selected for the final CTtLP battery, discarding 3 defective items from the initial 30-item draft.
Criterion-Related Validity
Criterion validity was established by evaluating the statistical association between students’ standardized CTtLP total scores and their ecological academic course performance in school. The correlation analysis revealed a statistically significant, moderate positive correlation:
r = 0.443 (p < 0.001)
This empirical coefficient confirms fair to strong criterion validity. It demonstrates that computational thinking, as measured by the CTtLP, is meaningfully tied to generalized academic achievement and cognitive performance in formal educational settings, while simultaneously maintaining sufficient divergent uniqueness to indicate that it measures a distinct cognitive capacity rather than duplicating conventional academic grades.
Construct Validity: Item Response Theory (3PL Model) Calibrations
To rigorously establish construct validity at the latent item level, Zhang and Wong (2023) calibrated the 27 items using a three-parameter logistic (3PL) Item Response Theory model, which estimates item discrimination (a), item difficulty (b), and pseudo-guessing (c) parameters:
- Item Discrimination (a-parameter): The test exhibited high discriminative power across all items, with a mean discrimination index of 2.188 (range: 1.437 to 3.331). Every item comfortably surpassed the conventional psychometric threshold of 0.80, confirming that the CTtLP items differentiate sharply between students with low versus high underlying computational thinking ability.
- Item Difficulty (b-parameter): The mean difficulty threshold across the 27 items was 0.353 (range: -1.096 to 1.892). This distribution demonstrates that the instrument spans an expansive continuum of cognitive challenge, providing sufficient low-difficulty baseline items to assess Grade 1 students accurately while incorporating high-difficulty items capable of measuring top-performing Grade 3 students without premature ceiling constraints.
- Pseudo-Guessing Parameter (c-parameter): In multiple-choice instruments targeted at young children, random guessing frequently introduces severe measurement noise. In the CTtLP, the mean guessing parameter was kept exceptionally low at 0.136 (range: 0.001 to 0.306). All items demonstrated guessing values well below the maximum theoretical threshold of 0.35, validating the effectiveness of integrating an explicit ‘I don’t know’ option to deter unreflective guessing.
8. Reliability
The reliability of the CTtLP was evaluated across three distinct psychometric dimensions: internal consistency reliability, temporal stability (test-retest reliability), and measurement precision along the latent continuum via the Item Response Theory Test Information Function (TIF).
Internal Consistency Reliability
The classical internal consistency of the CTtLP was quantified using Cronbach’s alpha ($lpha$). Across the normative validation sample of Grade 1 through 3 primary students, the scale achieved an internal consistency coefficient of:
$lpha$ = 0.873
In psychoeducational measurement, an alpha value exceeding 0.85 denotes high measurement reliability, confirming that the 27 items reliably measure a coherent, unified construct without excessive item redundancy or measurement error.
Test-Retest Temporal Stability
To evaluate whether computational thinking performance remains stable over time or fluctuates due to short-term developmental variance, temporal stability was examined across an eight-week interval. The Intraclass Correlation Coefficient was calculated:
ICC = 0.757
An ICC value of ~0.76 over an extensive two-month window provides clear evidence of temporal stability. This demonstrates that the CTtLP measures enduring cognitive-developmental competencies rather than transient situational state fluctuations.
Test Information Function (TIF) and Latent Precision
Modern psychometric theory recognizes that measurement reliability is not a uniform, static figure across all performance levels; rather, it varies across the latent ability continuum ($ heta$). Using the 3PL IRT model, the Test Information Function was mapped. The empirical data showed that the total test information peaked at an impressive value of 14.65 when participant ability was approximately $ heta = 0.9$.
This distribution reveals that the CTtLP offers maximum measurement precision for students exhibiting slightly above-average computational ability, while sustaining stable psychometric precision across the broader ability range (-1.5 < $ heta$ < 2.0). Consequently, standard errors of measurement remain minimal across the majority of the lower primary population, making the tool well-suited for screening both normative and gifted cohorts.
9. Factor Analysis
To assess the internal dimensional architecture of the 27-item instrument, Confirmatory Factor Analysis (CFA) was conducted to empirically test the theoretical assumption that the CTtLP operates as a unidimensional measurement model.
Confirmatory Factor Analysis (CFA) Fit Indices
A single-factor structural equation model—where all 27 items load onto a unified latent computational thinking factor—was evaluated against the primary validation dataset from public elementary schools in northern China. In accordance with benchmark criteria established by Hu and Bentler (1999), the single-factor model exhibited solid fit to the observed data:
- Root Mean Square Error of Approximation (RMSEA): 0.041 (Threshold < 0.06 indicates close, excellent fit; values below 0.05 represent exceptional model-data correspondence).
- Comparative Fit Index (CFI): 0.911 (Exceeds the conventional baseline threshold of > 0.90, confirming acceptable relative fit against a null independence model).
- Tucker-Lewis Index (TLI): 0.900 (Meets the standard psychometric cutoff of ≥ 0.90, verifying appropriate model parsimony).
These fit indices empirically corroborate the structural unidimensionality of the CTtLP. While the underlying tasks involve varied contextual themes (e.g., executing instructions, identifying errors, tracking path trajectories across different graphic scenarios), they load uniformly onto a single dominant latent cognitive trait: lower primary computational thinking ability. This unidimensionality satisfies the essential local independence assumptions required for valid IRT parameter estimation and permits psychometricians to compute a single aggregated composite score representing the child’s overall computational reasoning capability.
10. Instrument / Measurement Tool
The administrative structure, formatting properties, and scoring specifications of the Computational Thinking Test for Lower Primary are detailed below:
- Test Type: Standardized psychoeducational cognitive achievement measure; constructed-response multiple-choice inventory.
- Target Population: Lower primary school children (Grades 1 to 3, chronological ages 6 to 10 years). Validated across male and female students in public elementary school systems.
- Languages Available: English; Simplified Chinese.
- Total Item Count: 27 items (culled from an original 30-item developmental inventory through empirical psychometric screening).
- Test Scenarios: The 27 items are distributed systematically across three real-world, child-friendly narrative scenarios:
- Scenario A: Path Traversal and Instruction Sequence Tracking (mentally tracing concrete movements across grid networks).
- Scenario B: Algorithmic Transformation and Output Prediction (determining the end-state of objects following multi-step operational chains).
- Scenario C: Algorithmic Debugging and Error Identification (detecting the exact step, orientation, or instruction that prevents the attainment of a stated objective).
- Response Format: Multiple-choice format containing five (5) discrete options per item:
- One (1) keyed correct target response.
- Three (3) theoretically derived distractor options reflecting common cognitive developmental misconceptions (e.g., incorrect frame of reference, premature sequencing termination, misinterpreting left/right rotational direction).
- One (1) explicit meta-cognitive option: “I don’t know”, included to lower guessing frequency and reduce pseudo-guessing distortion.
- Administration Time and Setting: Exactly 60 minutes are allocated for standard administration. The test can be administered in classroom group settings or individual testing rooms under the guidance of a trained proctor or educational psychologist.
- Scoring Protocol:
- Dichotomous Classical Scoring: Keyed correct response = 1 point; Incorrect distractor = 0 points; “I don’t know” response = 0 points; Omitted/blank = 0 points. Total raw score ranges from 0 to 27 points.
- Latent Trait Scoring (IRT): In research contexts, pattern scoring or Expected A Posteriori (EAP) / Maximum Likelihood Estimation (MLE) under the calibrated 3PL item parameters can be utilized to generate latent ability estimates ($ heta$) for each child.
11. Permissions & Fee and Test Year
- Year of Initial Publication: 2023.
- Copyright & Intellectual Property: © 2023 Shuhan Zhang and Gary K. W. Wong. All rights reserved under relevant educational copyright legislation.
- Commercial Fee: No commercial fee. The CTtLP was established as an academic, open-access diagnostic instrument designed to foster non-commercial empirical research and educational benchmarking.
- Permissions and Access Protocol: Qualified researchers, school administrators, and psychoeducational diagnosticians wishing to use the complete 27-item test forms, administrative user manuals, and graphical testing booklets must formally contact the corresponding author, Dr. Shuhan Zhang, via email at [email protected] or directly through the Faculty of Education at the University of Hong Kong.
12. References
- Aesaert, K., Voogt, J., Dochy, F., & van Braak, J. (2014). Design and validation of a computer-based assessment for digital literacy in lower secondary education. Computers & Education, 79, 1–14. https://doi.org/10.1016/j.compedu.2014.07.001
- Hu, L. T., & Bentler, P. M. (1999). Cutoff criteria for fit indexes in covariance structure analysis: Conventional criteria versus new alternatives. Structural Equation Modeling: A Multidisciplinary Journal, 6(1), 1–55. https://doi.org/10.1080/10705519909540118
- Mislevy, R. J., Steinberg, L. S., & Almond, R. G. (2003). On the structure of educational assessments. Measurement: Interdisciplinary Research and Perspectives, 1(1), 3–62. https://doi.org/10.1207/S15366359MEA0101_02
- Papert, S. (1980). Mindstorms: Children, computers, and powerful ideas. Basic Books, Inc.
- Piaget, J. (1952). The origins of intelligence in children. International Universities Press. https://doi.org/10.1037/11494-000
- Wing, J. M. (2006). Computational thinking. Communications of the ACM, 49(3), 33–35. https://doi.org/10.1145/1118178.1118215
- Zhang, S., & Wong, G. K. W. (2023). Development and validation of a computational thinking test for lower primary school students. Educational Technology Research and Development, 71(4), 1595–1630. https://doi.org/10.1007/s11423-023-10231-2