Clinical AssessmentDevelopmental PsychologyPsychometrics

Checklist for Autism in Toddlers- CHAT

A psychometric review of the Checklist for Autism in Toddlers (CHAT), examining its theoretical basis in Theory of Mind, validity, reliability, and administration rules.

memjavad
PUBLISHED
Scientifically Reviewed · Dr. Marwa Abd-Alazim · September 16, 2026
Medically & Scientifically Reviewed Verified: September 16, 2026
Dr. Marwa Abd-Alazim Ph.D.
Professor of Psychology University of Kerbala
Review Criteria & Clinical Standards

This content undergoes rigorous scientific peer-review and medical editorial standards at Arab Psychology Network to ensure clinical accuracy, validity, and compliance with evidence-based guidelines from leading psychological and healthcare authorities (APA / WHO).

Abstract

The Checklist for Autism in Toddlers (CHAT) is a seminal first-stage developmental screening instrument designed to identify children at risk for autism spectrum disorder (ASD) at 18 months of age. Developed in the United Kingdom by Simon Baron-Cohen, John Allen, and Christopher Gillberg (1992) and validated in extensive epidemiological cohorts by Gillian Baird and colleagues (2000), the CHAT operationalizes early sociocognitive deficits that manifest prior to full phenotypic language and behavioral presentations. The instrument consists of two distinct components totaling 14 items: Section A (9 parent-report items evaluating everyday behavioral patterns, joint attention, and pretend play) and Section B (5 direct observational items administered by a primary care professional, such as a general practitioner or health visitor). The psychometric architecture of the CHAT centers on five critical markers: protodeclarative pointing, gaze monitoring, and pretend play across both informant modalities.

Epidemiological evaluations spanning over 16,000 toddlers demonstrate that the CHAT exhibits exceptional specificity (99.9%) and a high positive predictive value (PPV; ~76% to 83% for broader developmental and spectrum disorders in high-risk profiles). However, population-based longitudinal tracking has revealed low sensitivity (approximately 18% to 38% for the entire spectrum), as toddlers with less severe presentations, higher intellectual ability, or developmental regressions often pass the 18-month threshold. Despite these sensitivity limitations—which spurred the subsequent creation of modified instruments such as the M-CHAT and Q-CHAT—the original CHAT remains a foundational milestone in developmental psychometrics, establishing empirical evidence that neurodevelopmental sociocommunicative impairments can be systematically detected in infancy via targeted behavioral observation and parent report.

Keywords

Checklist for Autism in Toddlers, CHAT, Autism Spectrum Disorder, early screening, developmental psychometrics, joint attention, protodeclarative pointing, pretend play, Theory of Mind, infant assessment

Authors

The Checklist for Autism in Toddlers was conceived, refined, and validated by a consortium of developmental psychologists, child psychiatrists, and pediatricians across the United Kingdom and Sweden:

  • Sir Simon Baron-Cohen, PhD, FBA, FMedSci: Professor of Developmental Psychopathology, Department of Psychiatry, University of Cambridge; Director of the Autism Research Centre (ARC), Cambridge, United Kingdom.
  • John Allen, MB, ChB, FRCPsych: Consultant Child and Adolescent Psychiatrist, Child and Family Mental Health Services, Somerset NHS Trust, United Kingdom.
  • Christopher Gillberg, MD, PhD: Professor of Child and Adolescent Psychiatry, Gillberg Neuropsychiatry Centre, Sahlgrenska Academy, University of Gothenburg, Sweden.
  • Gillian Baird, FRCPCH: Consultant Paediatrician and Professor of Paediatric Neurodisability, Guy’s and St Thomas’ NHS Foundation Trust and King’s College London, United Kingdom (lead investigator for large-scale epidemiological trials).
  • Tony Charman, PhD: Professor of Clinical Child Psychology, Department of Psychology, Institute of Psychiatry, Psychology & Neuroscience (IoPPN), King’s College London, United Kingdom.
  • Antony Cox, FRCP, FRCPsych: Emeritus Professor of Child and Adolescent Psychiatry, Guy’s, King’s and St Thomas’ School of Medicine, London, United Kingdom.
  • John Swettenham, PhD: Division of Psychology and Language Sciences, University College London (UCL), United Kingdom.
  • Sally Wheelwright, MSc: Autism Research Centre, Department of Psychiatry, University of Cambridge, United Kingdom.
  • Ann Drew, PhD: Department of Child and Adolescent Psychiatry, Guy’s Hospital, King’s College London, United Kingdom.

Purpose

The primary clinical purpose of the CHAT is the universal and targeted secondary screening of 18-month-old toddlers to detect core behavioral markers indicative of autism and pervasive developmental disorders during primary care surveillance. Historically, autism was rarely diagnosed before 3 to 4 years of age, often because clinical evaluations awaited parental awareness of severe expressive language delay or obvious stereotyped behaviors. The CHAT was engineered to lower this diagnostic age threshold by focusing directly on early psychological capacities that emerge between 9 and 14 months in neurotypical development, specifically triadic joint attention and symbolic imagination.

From a public health and clinical surveillance perspective, the CHAT bridges parental observation with direct standardized clinical probing. Many parents may normalize early social deficits or lack comparison benchmarks; conversely, brief clinical encounters can lead to false impressions if a child is anxious, fatigued, or uncooperative. By pairing nine retrospective parent questions with five standardized, playful interactions administered during routine 18-month child health surveillance examinations (conducted primarily by Health Visitors or General Practitioners in the UK National Health Service), the instrument provides a structured, objective, and reproducible measurement protocol.

In research contexts, the CHAT provides an empirically validated paradigm to evaluate developmental trajectories, test sociocognitive models of early childhood development, and recruit prospective cohorts of infants displaying high-risk markers. Early identification facilitates timely referral to comprehensive multi-disciplinary diagnostic evaluations (e.g., Autism Diagnostic Observation Schedule [ADOS], Autism Diagnostic Interview-Revised [ADI-R]) and enrollment in evidence-based early behavioral interventions (such as the Early Start Denver Model or Naturalistic Developmental Behavioral Interventions). Initiating intervention during periods of maximal neuroplasticity prior to age 3 has been shown to produce substantial gains in functional communication, adaptive functioning, and cognitive outcomes.

Psychological Construct

The CHAT measures early sociocommunicative and imaginative competence, specifically indexing the presence, absence, or disruption of behaviors essential for normal sociocognitive development. The instrument does not measure broad motor development or general intelligence; rather, it isolates three foundational constructs: joint attention, symbolic or pretend play, and basic social orienting/engagement.

1. Joint Attention and Triadic Communication

Joint attention refers to the shared focus of two individuals on an external object or event, coordinated through gaze shifting, pointing, or vocalizations. The CHAT operationalizes this construct across two distinct behavioral mechanisms:

  • Protodeclarative Pointing (Items A7, Biv): The child uses an extended index finger not to request an object (which is protoimperative pointing, evaluated in Item A6), but purely to direct another person’s attention to an interesting target for social sharing (e.g., pointing at an airplane in the sky or a room light to elicit shared enjoyment). Protodeclarative pointing requires triadic coordination (child – adult – object) and indexes the psychological understanding that another person has their own internal attentional state that can be influenced.
  • Gaze Following / Gaze Monitoring (Item Bii): The ability of the child to track the line of sight and pointing gesture of an examiner across the room toward an ambiguous or distal object. This demonstrates an understanding of the referential nature of gaze and attention.

2. Symbolic and Pretend Play

Pretending represents the emergence of symbolic manipulation and mental representation (Items A5, Biii). In the CHAT, pretend play is assessed through simple sociodramatic scripts, such as pouring imaginary tea from a toy teapot into a cup and pretending to drink it or feed a doll. This capacity requires what developmental cognitive models term mental decoupling: the child must separate the primary perceptual identity of an object (e.g., a plastic cylinder) from its pretend identity (e.g., a warm cup of tea) while maintaining playful social engagement with an adult.

3. Social Interaction, Direct Engagement, and Imitation

Surrounding the primary sociocognitive axes are baseline social orienting behaviors: mutual eye gaze (Item Bi), social smiling, enjoyment of rough-and-tumble or vestibular social games (Item A1), social interest in other children (Item A2), interactive reciprocal games such as peek-a-boo (Item A4), and showing behaviors (Item A9: bringing objects to an adult to share interest rather than requesting help). Control items measuring functional sensorimotor play (Item A8), climbing (Item A3), and fine motor stacking (Item Bv) provide discriminant baseline data to distinguish isolated sociocommunicative deficits from global motor delays.

Theoretical Framework

The design of the CHAT is rooted in the cognitive-developmental Theory of Mind (ToM) framework formulated by Simon Baron-Cohen, Alan M. Leslie, and Uta Frith. ToM posits that human social navigation depends on the capacity to impute unobservable mental states—such as beliefs, desires, knowledge, intentions, and pretenses—to oneself and others to predict and interpret behavior.

Baron-Cohen conceptualized a multi-stage modular model of the developing mindreading system:

  • The Intentionality Detector (ID): A perceptual mechanism that interprets primitive motion stimuli in terms of basic volitional goals and desires (emerging in early infancy).
  • The Eye Direction Detector (EDD): A specialized system that detects the presence of eyes or eye-like stimuli, computes whether gaze is directed toward the self or elsewhere, and infers that the agent sees the target (emerging between 0 and 9 months).
  • The Shared Attention Mechanism (SAM): Emerging between 9 and 14 months, SAM builds triadic representations by linking the self, an agent, and an object (e.g., [Self sees that Agent sees Object]). SAM integrates perceptual outputs from ID and EDD, providing the cognitive architecture required for protodeclarative pointing and gaze following.
  • The Theory of Mind Mechanism (ToMM): Typically maturing between 3 and 4 years, ToMM allows the processing of epistemic mental states, such as false beliefs and representational decoupling.

Concurrently, Alan M. Leslie’s (1987) representational model of pretend play posited that pretending necessitates a structural cognitive mechanism capable of decoupling an anchored primary representation of reality from a secondary pretend representation (e.g., “this banana is a telephone”), preventing representational abuse while preserving truth conditions about the physical world. Leslie argued that this decoupling engine forms the structural precursor to attributing mental states (such as belief) to others.

Baron-Cohen and colleagues hypothesized that the pathognomonic neurocognitive marker of autism is a primary congenital disruption or developmental failure of the Shared Attention Mechanism (SAM) and the pretend play decoupling mechanism. Consequently, while general intelligence, rote memory, and protoimperative gestures (pointing to demand an object, which only requires understanding agency rather than shared mental states) may remain intact, triadic joint attention and symbolic pretend play will be selectively absent at 18 months. The CHAT was engineered to directly evaluate the behavioral manifestations of this theoretical deficit.

Validity

The psychometric validity of the CHAT has been evaluated through clinical validation studies, large-scale general population trials, and prospective longitudinal cohort follow-ups.

1. Construct and Content Validity

Content validity was established by selecting behaviors that developmental psychopathology and experimental psychology had demonstrated to be universally present in neurotypical 18-month-old children but selectively absent or impaired in young autistic children. In initial clinical validation trials (Baron-Cohen et al., 1992, 1996), the CHAT was administered to high-risk siblings of autistic probands, neurotypical infants, and children with non-autistic developmental delays. The key items—specifically protodeclarative pointing (A7, Biv), gaze following (Bii), and pretend play (A5, Biii)—consistently segregated children who subsequently received a diagnosis of autism from both neurotypical peers and developmentally delayed controls, confirming high construct validity.

2. Predictive and Discriminant Validity (The South Thames Population Study)

The primary benchmark for CHAT validation was a large-scale, prospective, population-based epidemiological study conducted in the South Thames region of England (Baird et al., 2000; Baron-Cohen et al., 2000). A total of 16,235 children were screened by Health Visitors and GPs at 18 months of age, with comprehensive multidisciplinary clinical follow-up extending through age 7. Children were classified into stratified risk tiers based on their response patterns:

  • High Risk for Autism: Failing all five key markers (parent report of pretend play [A5] and pointing for interest [A7], combined with clinician observation of gaze following [Bii], pretend play [Biii], and pointing for interest [Biv]). In the South Thames cohort, 38 children were identified as High Risk. Longitudinal clinical diagnostic evaluation revealed that 83% of these children were diagnosed with an autism spectrum disorder or severe developmental disability, demonstrating excellent Positive Predictive Value (PPV) for this specific failure profile.
  • Medium Risk for Autism: Failing pointing for interest (A7, Biv) without necessarily failing pretend play. This group identified toddlers at risk for broader communication impairments, social-communication deficits, or less classical presentations.
  • Discriminant Power: The specificity of the CHAT in this population trial was exceptionally high: 99.9% (95% CI [99.8%, 100.0%]). The tool rarely produced false positives among normally developing children or those with isolated motor delays. The false-positive cases primarily comprised children with severe non-spectrum global cognitive delays or severe receptive language impairments.

3. Sensitivity Limitations

While specificity and PPV for the “High Risk” profile were outstanding, the population sensitivity of the CHAT was low. At age 7 follow-up, epidemiological tracking revealed that the CHAT had identified only approximately 18% to 38% of all children who ultimately met diagnostic criteria for ASD within the birth cohort. Toddlers who passed the CHAT at 18 months but later received an ASD diagnosis frequently exhibited higher IQ, developed late-emerging deficits, experienced regression after 18 months, or displayed subtler social impairments (e.g., Asperger syndrome or PDD-NOS). These empirical findings led to two conclusions: while a positive (failed) CHAT result carries significant clinical validity, a negative (passed) CHAT result cannot rule out subsequent emergence of an autism spectrum disorder.

Reliability

Evaluating the reliability of infant developmental screening tests requires examining inter-rater agreement across observers and temporal stability across short screening windows.

1. Inter-Rater Reliability

Inter-rater reliability is particularly critical for Section B, where a clinician must elicit and judge child behaviors under standardized conditions. In the validation cohorts reported by Baron-Cohen et al. (1996) and Charman et al. (1997, 1998), two independent clinicians concurrently observed and scored toddler responses during administration of Section B. Agreement on the critical items was remarkably high:

  • Item Bi (Eye contact): κ = 0.82 to 0.88
  • Item Bii (Gaze following): κ = 0.86 to 0.94
  • Item Biii (Pretend play): κ = 0.84 to 0.91
  • Item Biv (Pointing to indicate interest): κ = 0.88 to 0.95
  • Item Bv (Brick stacking): κ = 0.92 to 0.98

Overall inter-rater concordance for classifying a child as passing versus failing Section B consistently exceeded 90% (κ > 0.85), indicating that structured scoring guidelines and behavioral prompts minimize subjective clinical drift.

2. Test-Retest Reliability and Rescreening Protocols

Toddler behavior in clinical environments is sensitive to acute fatigue, illness, transient separation anxiety, and uncooperativeness. To avoid transient situational false positives, the CHAT protocol incorporates a mandatory two-stage re-screening rule. Children who failed the critical items at the initial 18-month visit were re-assessed approximately one month later (at 19 months) using the identical protocol. In the Baird et al. (2000) trial, a substantial proportion of children who initially failed Section B due to transient non-compliance passed upon re-testing. Stability of failure across the one-month test-retest interval approached 100% true-positive status for developmental pathology. Thus, while single-administration test-retest stability may reflect state-dependent infant factors, sequential two-stage screening confers robust temporal stability to the classification decision.

Factor Analysis & Structural Characteristics

Because the CHAT is a short, categorical, non-continuous screening checklist composed of binary (Yes/No) classifications, standard linear factor analyses (e.g., Pearson-based exploratory factor analysis) are psychometrically inappropriate. Methodologists have evaluated its internal structure using multidimensional scaling, latent class analysis, and non-parametric item analysis for binary variables.

1. Structural Item Clustering

Item intercorrelation matrices and categorical latent variable models from population-based screening data reveal a distinct two-cluster architecture:

  • Cluster 1: Sociocommunicative Joint Attention and Symbolic Competence: Formed by the five critical items (A5, A7, Bii, Biii, Biv). These items exhibit strong mutual conditional dependence. Latent class models demonstrate that performance on this cluster differentiates children along a severe social-communication dimension with high latent factor loadings (> 0.75). A failure across this cluster rarely occurs in isolation from significant neurodevelopmental disruption.
  • Cluster 2: General Motor, Reciprocal Play, and Instrumental Behavior: Formed by items assessing protoimperative pointing (A6), brick stacking (Bv), climbing (A3), rough-and-tumble play (A1), and handling small toys (A8). These items load onto a distinct general functional maturity dimension. Factor analytic models demonstrate that children who fail Cluster 1 often pass Cluster 2, confirming structural divergence between social-cognitive signaling and general sensorimotor development.

2. Item Response Characteristics and Item Difficulty

Within an Item Response Theory (IRT) framework for dichotomous screening items, the critical CHAT items function as high-difficulty, high-discrimination thresholds. In normal populations at 18 months, item endorsement (“Yes”) is nearly at ceiling (> 97% pass rate for each individual item). Consequently, item difficulty parameters (β) are positioned far into the negative latent ability spectrum (where ability = social communication competence), meaning that failing an item reflects a severe deviation from normal developmental timing. The discrimination parameters (α) for Items Biv (observation of pointing to show interest) and Biii (elicited pretend play) are exceptionally high, confirming their role as decisive anchors in the decision algorithm.

Instrument / Measurement Tool

The Checklist for Autism in Toddlers is a 14-item, multi-informant screening tool completed in approximately 5 to 10 minutes within a primary care or clinical research setting.

1. Structure and Administration Format

  • Test Type: Multi-informant developmental screening instrument (Parent Questionnaire + Direct Standardized Behavioral Observation).
  • Target Population: Toddlers aged 18 months (usable clinically between 18 and 24 months).
  • Total Item Count: 14 items.
    • Section A: 9 parent-report items administered verbally or self-completed by the parent/caregiver.
    • Section B: 5 observational items administered directly by the health professional (GP, Health Visitor, or Pediatrician).
  • Response Scale: Binary categorical options: YES or NO.

2. Standardized Behavioral Probing Rules (Section B)

Section B requires specific physical stimuli (e.g., a miniature toy cup and teapot, small building blocks/bricks, and distal target objects in the examination room):

  • Item Bii (Gaze Monitoring): The clinician must secure the child’s attention, point across the room to an engaging object while exclaiming, “Oh look! There’s a (name of toy)!”, and observe the child’s face. To score YES, the child must shift gaze away from the clinician’s hand and directly inspect the distant object.
  • Item Biii (Pretend Play): The clinician presents a toy teapot and cup, saying “Can you make a cup of tea?” To score YES, the child must spontaneously or upon suggestion act out a pretend sequence (e.g., pouring tea, sipping, offering the cup). Any other elicited pretend sequence (e.g., feeding a teddy bear) also qualifies.
  • Item Biv (Protodeclarative Pointing): The clinician asks, “Where’s the light?” or “Show me the light” (or an equivalent unreachable target like a teddy bear on a high shelf). To score YES, the child must point with an extended index finger toward the target and simultaneously or immediately check back to make eye contact with the clinician’s face.
  • Item Bv (Brick Stacking): The child is invited to build a tower of blocks to establish fine motor compliance and baseline cooperative responsiveness.

3. Scoring and Risk Classification Rules

Scoring does not use a simple additive sum score; instead, it relies on a qualitative pattern-matching decision algorithm focused on five critical items:

  • Item A5: Parent report of pretend play
  • Item A7: Parent report of protodeclarative pointing (pointing to show interest)
  • Item Bii: Clinician observation of gaze following
  • Item Biii: Clinician observation of pretend play
  • Item Biv: Clinician observation of protodeclarative pointing
Risk Category Algorithmic Criteria Clinical Action
High Risk for Autism Failing (scoring NO on) all 5 critical items: A5, A7, Bii, Biii, and Biv. Re-screen after 4 weeks. If failure pattern persists, immediate direct referral to specialist developmental pediatric/psychiatric services.
Medium Risk for Autism Failing pointing to indicate interest on both parent report and clinical observation (A7 and Biv), without necessarily failing pretend play or gaze following. Re-screen after 4 weeks. If sustained, refer for developmental language and communication assessment.
Low / Negative Risk Passing all or most critical markers; failure on non-critical items only (e.g., A1, A3, A8). Standard primary care developmental surveillance.

Permissions & Fee and Test Year

The original Checklist for Autism in Toddlers was developed and published in 1992 by Simon Baron-Cohen, John Allen, and Christopher Gillberg, followed by the large-scale population validation protocols published in 1996 and 2000.

The CHAT was placed in the public domain for clinical, educational, and non-commercial scientific research purposes to advance the early detection of autism spectrum disorders globally. No licensing fees or royalty payments are required to administer the instrument. The instrument and its accompanying scoring guides can be accessed via academic publications and through the Autism Research Centre (ARC) at the University of Cambridge.

Clinicians and investigators must retain the standardized wording, observational guidelines, and scoring algorithms intact when utilizing the instrument to ensure psychometric validity. Commercial reproduction, incorporation into closed-source diagnostic software suites, or for-profit dissemination requires explicit formal permission from the original copyright holders and authors.

References

  • Baird, G., Charman, T., Baron-Cohen, S., Cox, A., Swettenham, J., Wheelwright, S., & Drew, A. (2000). A screening instrument for autism at 18 months of age: A 6-year follow-up study. Journal of the American Academy of Child & Adolescent Psychiatry, 39(6), 694–702. https://doi.org/10.1097/00004583-200006000-00007
  • Baird, G., Charman, T., Cox, A., Baron-Cohen, S., Swettenham, J., Wheelwright, S., & Drew, A. (2001). Screening and surveillance for autism and pervasive developmental disorders. Archives of Disease in Childhood, 84(6), 468–475. https://doi.org/10.1136/adc.84.6.468
  • Baron-Cohen, S., Allen, J., & Gillberg, C. (1992). Can autism be detected at 18 months? The needle, the haystack, and the CHAT. The British Journal of Psychiatry, 161(6), 839–843. https://doi.org/10.1192/bjp.161.6.839
  • Baron-Cohen, S., Cox, A., Baird, G., Swettenham, J., Nightingale, N., Morgan, K., Drew, A., & Charman, T. (1996). Psychological markers in the detection of autism in infancy in a large population. The British Journal of Psychiatry, 168(2), 158–163. https://doi.org/10.1192/bjp.168.2.158
  • Baron-Cohen, S., Wheelwright, S., Cox, A., Baird, G., Charman, T., Swettenham, J., Drew, A., & Doehring, P. (2000). Early identification of autism by the Checklist for Autism in Toddlers (CHAT). Journal of the Royal Society of Medicine, 93(10), 521–525. https://doi.org/10.1177/014107680009301007
  • Charman, T., Swettenham, J., Baron-Cohen, S., Cox, A., Baird, G., & Drew, A. (1997). Infants with autism show atypical fronto-striatal brain function: A longitudinal social-communication study. Science, 278(5336), 302–305. https://doi.org/10.1126/science.278.5336.302
  • Leslie, A. M. (1987). Pretense and representation: The origins of “theory of mind.” Psychological Review, 94(4), 412–426. https://doi.org/10.1037/0033-295X.94.4.412
  • Robins, D. L., Fein, D., Barton, M. L., & Green, J. A. (2001). The Modified Checklist for Autism in Toddlers: An initial study investigating the early detection of autism and pervasive developmental disorders. Journal of Autism and Developmental Disorders, 31(2), 131–144. https://doi.org/10.1023/A:1010738829569

13. Items of the Scale (Questionnaire)

Below are the authentic scale items in their original language as published in the standard psychometric validation studies, without modification or translation to preserve instrument validity and reliability:
1

Does your child enjoy being swung‚ bounced on your knee‚ etc?
2

Does your child take an interest in other children?
3

Does your child like climbing on things‚ such as up stairs?
4

Does your child enjoy playing peek-a-boo/hide-and-seek?
5

Does your child ever PRETEND‚ for example‚ to make a cup of tea using a toy cup and teapot‚ or pretend other things?
6

Does your child ever use his/her index finger to point‚ to ASK for something?
7

Does your child ever use his/her index finger to point‚ to indicate INTEREST in something?
8

Can your child play properly with small toys (eg. cars or bricks) without just mouthing‚ fiddling or dr‎opping them?
9

Does your child ever bring objects over to you (parent) to SHOW you something?

Rate This Scale

5.0 / 5 1 vote

Cite This Article

memjavad (2026, September 16). Checklist for Autism in Toddlers- CHAT. PSYCHOLOGICAL DATABASE. https://en.arabpsychology.com/scales/checklist-for-autism-in-toddlers-chat/
memjavad. “Checklist for Autism in Toddlers- CHAT.” PSYCHOLOGICAL DATABASE, 16 September 2026, https://en.arabpsychology.com/scales/checklist-for-autism-in-toddlers-chat/.
memjavad. “Checklist for Autism in Toddlers- CHAT.” PSYCHOLOGICAL DATABASE. September 16, 2026. https://en.arabpsychology.com/scales/checklist-for-autism-in-toddlers-chat/.