Cognitive ScienceDevelopmental PsychologyTheory of Mind

The Smarties Task (False Belief) – Josef Perner, Susan Leekam, and Heinz Wimmer

A comprehensive academic analysis of the Smarties Task, developed by Josef Perner, Susan Leekam, and Heinz Wimmer to assess theory of mind and false belief.

memjavad
PUBLISHED
Scientifically Reviewed · Dr. Marwa Abd-Alazim · September 12, 2026
Medically & Scientifically Reviewed Verified: September 12, 2026
Dr. Marwa Abd-Alazim Ph.D.
Professor of Psychology University of Kerbala
Review Criteria & Clinical Standards

This content undergoes rigorous scientific peer-review and medical editorial standards at Arab Psychology Network to ensure clinical accuracy, validity, and compliance with evidence-based guidelines from leading psychological and healthcare authorities (APA / WHO).

The philosophical investigation of how the human mind grasps the interiority of other minds represents one of the most enduring inquiries in cognitive science. For centuries, epistemologists grappled with the problem of other minds through introspective, rationalist, or strictly behavioral lenses. However, the empirical revolution in developmental psychology during the late twentieth century dramatically transformed these abstract conjectures into rigorous experimental science. Central to this paradigm shift was the formulation of Theory of Mind (ToM)—the neurocognitive capacity to attribute unobservable mental states, such as beliefs, desires, intentions, and knowledge, to oneself and others, and to understand that such internal representations systematically govern human action, even when they diverge from physical reality.

While early developmental investigations demonstrated that young children rapidly acquire an understanding of tangible desires and perceptual perspectives, the ultimate benchmark for a fully representational Theory of Mind became the false-belief task. To comprehend a false belief is to grasp epistemic opacity: the realization that internal representations do not merely mirror reality directly, but instead serve as mediated interpretations that can be objectively incorrect, outdated, or outright fabricated. Without this cognitive architecture, the social world appears as a bewildering sequence of unmediated mechanical reactions, stripping human behavior of its narrative coherence, deceit, irony, and communicative subtlety.

Among the empirical paradigms engineered to probe this cognitive milestone, none has achieved greater methodological elegance or theoretical resonance than the unexpected contents paradigm, colloquially known worldwide as the Smarties Task. Designed and published in 1987 by developmental psychologists Josef Perner, Susan Leekam, and Heinz Wimmer, this experimental protocol simplified the taxing narrative and spatial demands of earlier location-change designs. By presenting children with an iconic, highly recognizable confectionery container concealing an unexpected object—such as pencils—the experiment elegantly isolated both first-person autobiographical belief revision and third-person prospective attribution. The resulting empirical findings not only illuminated the profound conceptual revolution occurring between three and five years of age in neurotypical development, but also provided a diagnostic instrument that exposed the specific socio-cognitive architecture of neurodevelopmental conditions such as Autism Spectrum Disorder.

1. Historical and Theoretical Foundations of Theory of Mind

1.1 Premack and Woodruff’s Foundational Inquiries

The empirical genesis of Theory of Mind research did not originate within pediatric laboratories, but rather in the field of comparative primatology. In their seminal 1978 paper, “Does the chimpanzee have a theory of mind?”, David Premack and Guy Woodruff sought to determine whether an adult chimpanzee named Sarah could interpret the deliberate actions of human actors by attributing internal motivational and epistemic states to them. Premack and Woodruff presented Sarah with videotaped sequences of human actors struggling to resolve practical quandaries, such as shivering near an unlit heater or attempting to retrieve out-of-reach bananas. Sarah consistently selected photographs depicting functional solutions—such as an uncoiled hose attached to a faucet or a lit match. The researchers posited that Sarah’s predictive capability relied upon imputing desires and knowledge states to the human actors, concluding that chimpanzees operate with an intentional framework.

This provocative conclusion immediately catalyzed intense debate across philosophy of mind and cognitive ethology. Philosophers Daniel Dennett and Jonathan Bennett published pivotal critiques that fundamentally reshaped the experimental criteria for demonstrating mentalizing capacities. Dennett, drawing upon his framework of the “intentional stance,” observed that Sarah’s successful selections could be explained through non-mentalistic heuristics, such as associative learning, physical contingency matching, or perceptual script fulfillment. An organism might deduce that a key pairs with a lock purely through behavioral conditioning, entirely devoid of any meta-representational awareness that the human actor holds an internal goal or believes the door to be locked.

Bennett and Dennett converged on a foundational epistemic principle: to unequivocally prove that an agent possesses a Theory of Mind, one must demonstrate that the agent can attribute an internal mental representation that actively contradicts reality. If the attributed belief is identical to current physical reality, an empirical observer cannot disentangle whether the subject is tracking the external state of the physical world or the internal state of another mind. Thus, the deliberate design of experimental protocols assessing false belief became the definitive litmus test for mentalistic attribution. Only when an organism predicts that another agent will act in accordance with an objectively erroneous mental model can associative explanations be discarded in favor of genuine representational attribution.

1.2 The Evolution from Location Change to Deceptive Containers

In response to the philosophical gauntlet thrown down by Dennett and Bennett, Austrian developmental psychologists Heinz Wimmer and Josef Perner engineered the first empirical false-belief test for human ontogeny in their landmark 1983 study. Known historically as the “Maxi and the Chocolate” task—the intellectual progenitor of the popularized Sally-Anne paradigm—this unexpected-transfer protocol evaluated whether young children could track the epistemic displacement of an object. In the experimental narrative, a protagonist named Maxi places chocolate into a blue cupboard and leaves the kitchen to play. In his absence, his mother transfers the chocolate into a green cupboard. Maxi returns, and the child participant is posed the critical test question: “Where will Maxi look for his chocolate?”

Wimmer and Perner discovered that while five-year-olds correctly pointed to the blue cupboard—acknowledging Maxi’s out-of-date mental representation—children aged three, and many aged four, exhibited an unyielding “realist bias,” predicting that Maxi would look in the green cupboard where the chocolate physically resided. Despite its revolutionary impact, the unexpected-transfer paradigm suffered from considerable methodological friction. The protocol imposed substantial extraneous cognitive overhead: children had to parse an extended multi-act narrative, maintain the spatial coordinates of two identical cupboards in working memory, track a dynamic sequence of physical displacements, and inhibit the visual salience of the current location while calculating the narrative timeline.

These procedural burdens sparked contentious debates regarding whether young children genuinely lacked a representational Theory of Mind or simply succumbed to performance limitations, including working memory exhaustion, narrative decay, and spatial confusion. To eliminate these confounding variables, developmental researchers recognized the theoretical imperative for a static, immediate, and non-narrative stimulus. The objective was to design a paradigm that bypassed physical relocation altogether, shifting the diagnostic mechanism from spatial-tracking dynamics to a direct representational mismatch rooted in an object’s deceptive outward appearance versus its internal reality.

1.3 Perner, Leekam, and Wimmer’s 1987 Collaborative Innovation

To overcome the intrinsic limitations of the Maxi paradigm, Josef Perner, Susan Leekam, and Heinz Wimmer joined forces in a collaborative study published in 1987 under the title “Lack of belief about belief: Young children’s and autistic people’s failure to understand both the true and false belief of other people.” Working across institutions in the United Kingdom and Austria, their explicit intention was to isolate epistemic mentalizing from the confounding mechanical demands of tracking physical relocations across space. They sought to invent a protocol characterized by sensory immediacy, cultural resonance, and narrative minimalism.

The resulting innovation was the deceptive box, or unexpected contents paradigm, colloquially designated as the Smarties Task due to their utilization of the ubiquitous cylindrical cardboard tube manufactured by Rowntree’s. By employing packaging universally recognized by European children as containing distinct, sugar-coated chocolate lentils, the researchers bypassed the need to establish artificial narrative premises. The packaging itself activated a powerful, pre-existing conceptual schema in the child’s mind. When the physical container was opened to reveal an incongruous set of objects—such as ordinary writing pencils—the paradigm generated an immediate, tangible rupture between mental expectation and physical actuality.

Furthermore, Perner, Leekam, and Wimmer structured the task to capture a dimension of social cognition entirely inaccessible to the Maxi unexpected-transfer protocol: the systematic comparison between first-person (self) and third-person (other) epistemic states. Because the child was not an omniscient third-party observer watching puppets, but an active participant who was personally deceived by the packaging, the experimenters could evaluate how children process the temporal revision of their own false beliefs alongside their attributions to peers. The simultaneous deployment of this paradigm with neurotypical preschoolers and clinical cohorts marked an enduring paradigm shift in experimental cognitive development.

2. The Experimental Protocol of the Smarties Task

2.1 Phase 1: Initial Presentation and Inferred Expectations

The experimental protocol of the Smarties Task is celebrated for its stark, deceptive simplicity and rigorous procedural standardization. The session begins in a controlled, distraction-minimized environment where the experimenter sits directly across from the child participant. Without providing verbal preambles or leading narrative cues, the experimenter retrieves a standard commercial confectionery container—the iconic cylindrical Smarties tube, characterized by its vibrant background and colorful graphic representations of chocolate candies—and presents it within the child’s direct visual and tactile field.

Once the container is placed on the table, the experimenter initiates the baseline elicitation phase by posing the fundamental question: “What do you think is inside this box?” (or “What is in here?”). This prompt is intentionally framed using canonical epistemic verbs designed to tap into the child’s spontaneous semantic associations. The child does not observe an arbitrary neutral object; rather, the encounter stimulates deeply ingrained perceptual and cultural categorization schemas. Decades of market familiarity and socio-cultural exposure ensure that the child immediately maps the exterior semiotic markers—the commercial typography, brand branding, and structural iconography—onto internal category representations of confectionery.

Empirical integrity mandates that the experimenter meticulously documents this initial response. Across hundreds of replications worldwide, neurotypical children aged three and older overwhelmingly respond without hesitation: “Smarties,” “sweets,” or “candy.” If a child exhibits hesitation or provides an idiosyncratic response, standard testing protocols require the experimenter to clarify the package identity to ensure that a verified baseline expectation is established. Without this verified activation of an initial epistemic state, subsequent phases of the experiment lose diagnostic validity, as the cognitive surprise is entirely dependent upon the disruption of this initial schema.

2.2 Phase 2: The Deceptive Revelation and Reality Confirmation

Following the explicit vocal confirmation of the child’s expectation, the experiment transitions immediately to Phase 2: the deceptive revelation. The experimenter maintains neutral affect, refraining from theatrical gestures or linguistic hyperbole, and gently removes the plastic cap or slides open the container. Instead of the anticipated bright chocolate sweets cascading onto the testing surface, the experimenter withdraws the true physical contents: typically five small, sharpened wooden pencils, though some experimental variations employ matches, buttons, or small spoons.

To eliminate any perceptual ambiguity, sensory illusion, or conceptual vagueness, the experimenter encourages the child to engage directly with the revealed objects. The child is permitted to touch, handle, or manipulate the pencils, anchoring the true state of physical reality through multiple sensory modalities. The experimenter then issues an explicit reality control prompt: “Look, they are pencils! What are they really?” The child must verbally confirm the physical identity of the revealed object by stating, for example, “pencils.” This step secures empirical control over simple perceptual or vocabulary deficits; if a child cannot correctly identify or articulate the physical object sitting directly before them, the testing protocol is invalidated.

Once the child has physically interacted with the atypical contents and verbally validated their true identity, the experimenter gathers the pencils and meticulously places them back inside the original container. The cap is securely replaced, restoring the deceptive package to its original, unopened exterior appearance. The container remains situated on the table, visually identical to its state during Phase 1, yet conceptually transformed into a nexus of competing cognitive realities.

2.3 Phase 3: The Critical Representational Test Questions

Phase 3 represents the theoretical core of the experimental design, where the experimenter presents a battery of tightly controlled representational test questions designed to evaluate epistemic decoupling. The questions must be administered with strict syntactic precision to mitigate linguistic scaffolding or conversational interference. The protocol bifurcates into two distinct epistemological vectors: the retrospective evaluation of the self’s prior mental state, and the prospective evaluation of a naive other’s mental state.

The Self-Belief Test Question (evaluating autobiographical representational change) is articulated as: “When you first saw this box, before we opened it, what did you think was inside?” The question forces the child to retrieve an episodic mental state from the immediate past—a state that has since been proven factually wrong. To successfully answer, the child must refer not to what is currently inside the container, but to the counterfactual cognitive state they personally inhabited moments earlier.

The Other-Belief Test Question (evaluating third-person epistemic attribution) is framed around an absent peer or a canonical puppet: “[Name of peer/friend], who is waiting outside, hasn’t seen inside this box. When [Peer] comes in and sees this box all closed up like this, what will they think is inside before we open it?” Here, the child must ignore their own privileged knowledge of the container’s true contents and project a naive epistemic state onto another mind based exclusively on the deceptive external appearance of the box.

To ensure diagnostic fidelity, these test questions are counterbalanced and tethered to the mandatory Reality Control Question: “What is actually inside the box right now?” Scoring criteria define absolute competence versus realist failure through a strict binary metric:

  • Pass Criteria: The child successfully decouples internal mental representation from physical reality by answering “Smarties” (or “sweets”) to the self-belief question, “Smarties” to the other-belief question, and “pencils” to the reality control question.
  • Realist Failure (Contamination): The child succumbs to realist contamination, answering “pencils” to the self-belief question (asserting they always knew it contained pencils) and/or “pencils” to the other-belief question (projecting their own privileged knowledge onto the naive observer), despite correctly identifying “pencils” on the reality control prompt.

3. Cognitive Asymmetry: Self-Belief versus Other-Belief Attribution

3.1 Autobiographical Memory and Representational Overwrite

One of the most striking, counterintuitive phenomena observed during the administration of the Smarties Task is the pervasive failure of young children to accurately report their own prior false beliefs. When three-year-old children are asked the self-belief question—“When you first saw this box, what did you think was inside?”—they routinely, and with absolute conviction, respond: “Pencils!” This dramatic failure cannot be attributed to standard episodic memory decay, as mere seconds have elapsed since the child enthusiastically declared that the tube contained candy.

This cognitive blind spot illustrates the psychological phenomenon of retroactive updating, or early developmental hindsight bias. In young children, current veridical feedback does not merely supplement prior cognitive states; it aggressively overwrites them. In a comprehensive empirical investigation comparing the unexpected contents task across diverse conditions, Alison Gopnik and Janet Astington (1988) demonstrated that children experience an acute inability to access their own historical epistemic trajectory once the physical reality is unveiled. The newly acquired knowledge acts as an unyielding cognitive attractor, collapsing the representation of the past into the reality of the present.

Gopnik and Astington posited that three-year-olds lack an internal, archival model of mind. Instead of viewing their minds as an interpretive historical log containing beliefs that evolve, update, or err, young children treat the mind as an unmediated cognitive camera. Once the lens captures the true state of affairs (pencils), the previous image is instantly eradicated. Consequently, claiming “I always knew it was pencils” is not an act of conscious mendacity, but an authentic manifestation of an autobiographical memory architecture that lacks the meta-representational hooks necessary to index superseded epistemic representations.

3.2 Decoupling Current Reality from Counterfactual Epistemic States

The cognitive operation underlying success in the Smarties Task requires what cognitive scientist Alan Leslie termed metarepresentation, and what Josef Perner defined as a fully realized “representational theory of mind.” The child must construct a primary representation of the physical world (a cardboard tube holding pencils) while simultaneously constructing and evaluating a secondary representation that points to that world (a mental state holding the proposition: “the tube contains sweets”).

This operation demands absolute cognitive decoupling. To credit a naive peer with the belief that the tube holds Smarties, the child must systematically isolate their own privileged, veridical knowledge within an epistemic quarantine. They must hold two mutually contradictory propositions in active cognitive suspension:

  • Proposition A (Ontological Reality): The container contains pencils.
  • Proposition B (Epistemic Representation): The subject/peer represents the container as holding Smarties.

Children under the age of four exhibit profound difficulty representing the complex relation between a representation and its referent. They cannot grasp epistemic opacity—the philosophical truth that an object can be represented under one description (a candy tube) while failing to satisfy that description in actuality. For the three-year-old, representations are transparent: they point straight through to reality. If the container holds pencils, then any thought about the container must inherently terminate in pencils. To project a false belief onto an observer requires an understanding that minds act as imperfect information processors, bound by the limits of perceptual access, rather than direct conduits to ontological truth.

3.3 Linguistic and Grammatical Modulation in False-Belief Questioning

The cognitive demands of the Smarties protocol are inextricably intertwined with linguistic competence. Psycholinguists have pointed out that the self-belief test question—“When you first saw this box, before we opened it, what did you think was inside?”—possesses a remarkably intricate syntactic and temporal architecture. It contains an embedded temporal adverbial clause (“When you first saw…”), a prepositional phrase specifying counterfactual conditions (“before we opened it”), and a complement-taking mental state verb (“think”) governing a subordinate proposition (“was inside”).

Subtle modulations in the phrasing of these prompts yield measurable fluctuations in performance. Research has revealed that highlighting temporal markers, such as explicitly stressing the temporal contrast—“What did you think was inside before we opened it?”—can provide pragmatic scaffolding that marginally elevates performance in transitional four-year-olds. However, for younger three-year-olds, syntactic scaffolding consistently fails to bridge the conceptual chasm. Even when the question is simplified to minimize memory strain, the realist response dominates.

Furthermore, experimental psychologists have evaluated whether children’s failures stem from conversational implicatures. For instance, children might interpret the question “What did you think was inside?” as an adult’s bizarre request for what the box *actually* contains, assuming the adult would not ask about an irrelevant past error. To resolve this semantic ambiguity, researchers deployed non-verbal, behavioral adaptations—such as having children select pictures representing their prior expectations or physically selecting items to prepare a trap for a peer. The persistence of false-belief failure across these non-verbal paradigms conclusively confirmed that the Smarties deficit is deeply conceptual, reflecting a genuine representational limitation rather than a mere semantic misunderstanding of adult syntactic constructions.

4. The Developmental Trajectory: The Critical Transition at Age Four

4.1 The Three-Year-Old Baseline and Realist Bias

Across vast developmental literature, the performance of neurotypical three-year-old children on the unexpected contents task is characterized by remarkable, near-universal empirical consistency: they fail. Regardless of variations in socioeconomic status, geographic location, or testing idiosyncrasies, children between 36 and 42 months demonstrate an overwhelming vulnerability to the realist bias. When confronted with the dual demands of self-revision and third-person attribution, the physical reality of the pencils entirely eclipses the ephemeral mental state of the past or of the naive peer.

This persistent failure illuminates the baseline cognitive architecture of the early preschool years. Three-year-olds are not devoid of social cognition; they possess an intuitive understanding of basic desires and perceptual gaze. They comprehend that an adult who wants an apple will reach for an apple, and they understand that a person whose eyes are closed cannot visually perceive an object. However, their social-cognitive architecture operates via what Josef Perner conceptualized as a “situation theory” or a “direct-copy theory” of mind. They view human behavior as tethered directly to states of the world, rather than to internal representations of those states.

Crucially, three-year-olds do not yet appreciate the causal etiology of knowledge. They fail to understand that perceptual access (looking, touching, hearing) is the mandatory prerequisite for epistemic formation. To a three-year-old, knowledge is not an acquired, internally constructed model; it is an objective attribute of the surrounding environment. Because the container physically houses pencils, the child assumes this information is universally accessible across all agents, rendering the attribution of an erroneous mental state logically impossible within their current cognitive framework.

4.2 The Four-Year-Old Conceptual Shift

Between 48 and 54 months of age, children undergo what developmental psychologists describe as one of the most profound cognitive transformations in early ontogeny: the transition from a direct-copy model of mind to a fully realized Representational Theory of Mind. Suddenly, the same child who months earlier adamantly claimed they always knew the tube contained pencils effortlessly bursts into laughter at the trick, readily answers “Smarties!” to the self-belief question, and giggles conspiratorially as they predict their friend will be completely fooled by the packaging.

The statistical robustness of this developmental watershed was definitively confirmed in a meta-analysis conducted by Henry Wellman, David Cross, and Julanne Watson (2001). Analyzing hundreds of false-belief studies encompassing thousands of individual participants across diverse paradigms, Wellman and colleagues demonstrated that age is the single most powerful predictor of false-belief performance. The developmental trajectory follows a steep, predictable logistic sigmoid curve: performance hovers near floor levels at 30 to 36 months, undergoes rapid acceleration between 42 and 48 months, and achieves consistent ceiling levels across neurotypical cohorts by 54 to 60 months.

This transition represents a foundational conceptual reorganization. The four-year-old child no longer explains human action solely by referencing physical affordances and raw reality. Instead, they incorporate an intermediary variable: the mental representation. They understand that human beings do not act upon the world as it objectively exists, but rather upon the world as they believe it to exist. Misrepresentation is recognized not merely as a sensory defect or an intentional lie, but as an inescapable property of subjective cognition. Consequently, the child unlocks advanced social capacities, including sophisticated pretend play, intentional concealment, irony, and empathetic perspective-taking.

4.3 Intermediate Stages: Partial and Inconsistent Success

Although the developmental shift appears categorically stark in cross-sectional comparisons, microgenetic and longitudinal studies expose a nuanced intermediate phase characterized by cognitive fluctuation, fragility, and context-dependent success. Children navigating this transitional window—typically between 40 and 46 months—frequently exhibit fragile competence that collapses under cognitive or environmental friction.

During this transitional phase, developmental scientists observe profound discrepancies between implicit and explicit markers of false-belief understanding. For instance, when eye-tracking mechanisms are paired with the Smarties Task, transitional children often display anticipatory looking behaviors: their visual gaze spontaneously fixates on the candy-related cues or exhibits prolonged pupil dilation (indicating cognitive surprise or expectation) when an experimenter acts upon the unexpected contents, even while their explicit verbal responses continue to succumb to the realist bias. This divergence suggests that implicit, sensorimotor representations of other minds precede the conscious, linguistic metarepresentations demanded by verbal questioning.

Furthermore, intermediate performance is highly susceptible to extraneous task demands. If the experiment introduces emotional distractions, increases working memory load, or utilizes less iconic packaging, the transitional child’s nascent false-belief competence rapidly fractures. The consolidation of representational coherence is not an instantaneous, modular switch, but a gradual process of cognitive stabilization, wherein representational principles are slowly abstracted from idiosyncratic perceptual experiences and generalized across diverse social environments.

5. Neurodevelopmental Applications: Autism Spectrum Disorder

5.1 Perner, Leekam, and Wimmer’s 1987 Empirical Investigation

The 1987 investigation by Josef Perner, Susan Leekam, and Heinz Wimmer holds monumental significance not merely for establishing the unexpected contents paradigm, but for pioneering its application to clinical populations. Working alongside the foundational inquiries initiated by Simon Baron-Cohen, Alan Leslie, and Uta Frith (1985) with the Sally-Anne task, Perner and colleagues sought to test whether the social and communicative impairments characteristic of autism stemmed from a specific neurocognitive deficit in Theory of Mind.

Perner, Leekam, and Wimmer administered the Smarties Task to three precisely calibrated cohorts:

  • Children diagnosed with Autism Spectrum Disorder (with chronological ages often exceeding 10 to 14 years, but with varying verbal mental ages);
  • Clinical control children with Down syndrome (chronologically older, but with intellectual disabilities matching the verbal mental age of the autistic cohort);
  • Young, neurotypical preschool children (aged three to four years).

The empirical findings revealed a stark double dissociation. The autistic individuals, despite possessing chronological and mental ages far higher than the neurotypical preschool baseline, demonstrated an extraordinary, persistent failure on the deceptive box task. In striking contrast, the Down syndrome cohort—despite having overall cognitive and intellectual limitations comparable to or lower than the autistic participants—passed both the self-belief and other-belief components of the Smarties Task at rates commensurate with neurotypical four- and five-year-olds.

The methodological elegance of the Smarties Task proved decisive in this discovery. Skeptics had argued that autistic children failed the earlier Sally-Anne paradigm due to its complex narrative structure, the tracking of doll personas, or difficulties in parsing spatial displacements. By reducing the experimental apparatus to an immediate, static, real-world confectionery container, Perner, Leekam, and Wimmer eliminated these narrative and spatial confounds, demonstrating that the deficit in autism is profoundly, uniquely social-epistemic.

5.2 The Mindblindness Hypothesis and Domain-Specific Cognitive Modules

The clinical findings of the 1987 Smarties study provided pivotal empirical fuel for Simon Baron-Cohen’s formulation of the Mindblindness Hypothesis. Baron-Cohen, Leslie, and Frith posited that the neurotypical human brain is genetically equipped with dedicated, domain-specific neurocognitive modules designed to process social agents. Within Alan Leslie’s model, the cognitive architecture matures through distinct hierarchical modules:

  • The Intentionality Detector (ID) and the Eye-Direction Detector (EDD), which emerge in early infancy to parse motion into primitive volitional goals and track visual orientation;
  • The Shared Attention Mechanism (SAM), developing between 9 and 14 months, which constructs triadic representations linking the self, an agent, and an external object;
  • The Theory of Mind Mechanism (ToMM), maturing around age four, which executes the metarepresentational operations necessary to decouple reality from belief, enabling false-belief computation.

According to this modular perspective, the selective failure of autistic individuals on the Smarties Task represents a profound, neurobiologically rooted impairment or developmental arrest of the Theory of Mind Mechanism (ToMM). The autistic brain can parse physical causality, mechanical transformations, and logical systems—often demonstrating average or superior performance on non-social spatial reasoning assessments—yet it exhibits selective blindness toward epistemic states. This modular dissociation provided compelling evidence against domain-general theories of human cognition, suggesting that navigating the physical universe of objects and navigating the social universe of intentional minds rely upon distinct neural and cognitive architectures.

5.3 Clinical Diagnostics and Translational Assessment Tools

The empirical robustness and clinical clarity of the unexpected contents paradigm led to its institutionalization as a core translational diagnostic assessment. Today, variations of the Smarties Task are embedded directly within gold-standard clinical instruments, most notably the Autism Diagnostic Observation Schedule (ADOS). In Module 3 and Module 4 of the ADOS—administered to verbally fluent children, adolescents, and adults—the unexpected contents protocol serves as an indispensable probe to evaluate meta-representational reasoning under naturalistic communicative pressures.

Beyond diagnostics, the mechanics of the Smarties Task have driven targeted therapeutic interventions. Cognitive-behavioral paradigms and socio-communicative training programs employ deceptive containers to visually and concretely externalize the invisible mechanics of the mind. By practicing with transparent versus opaque boxes, taking turns staging tricks, and utilizing visual thought bubbles, neurodivergent children are provided with explicit, compensatory cognitive scaffolds to deduce epistemic states that neurotypical children acquire through automatic, intuitive development.

Nonetheless, clinical developmentalists emphasize the boundaries of laboratory false-belief tasks. Passing an explicit, structured Smarties Task in a serene, one-on-one laboratory environment does not immediately translate to fluid, naturalistic social competence in chaotic real-world settings. Many high-functioning adolescents with autism learn to solve false-belief tasks via explicit, algorithmic reasoning—treating the problem as a deterministic logic puzzle—yet continue to experience profound exhaustion and confusion when required to navigate the rapid, implicit, and multi-layered mentalizing demands of everyday social discourse.

6. Cognitive Mechanisms: Executive Functions and Inhibitory Control

6.1 Inhibition of the Salient Realist Response

While domain-specific theorists attributed failure on the Smarties Task strictly to an absent or impaired mentalizing module, cognitive psychologists advanced an alternative, domain-general account centered upon the development of Executive Functions (EF). Central to this perspective is the concept of inhibitory control: the neurocognitive ability to suppress a prepotent, automatic behavioral response in favor of an internally maintained, goal-directed alternative.

When evaluated through this computational lens, the Smarties Task presents an exceptionally hostile inhibitory challenge. Consider the child’s cognitive state at Phase 3: the pencils have just been revealed, physically manipulated, and verified. The physical reality of the pencils is perceptually hyper-salient, cognitively active, and reinforced by immediate sensory feedback. When the experimenter poses the test question—“What did you think was inside?”—the neural representation of “pencils” fires with immense activation strength. To formulate the correct answer—”Smarties”—the child must deploy prefrontal inhibitory mechanisms to actively suppress the prepotent impulse to name the physical reality residing right before their eyes.

Pioneering empirical work by Stephanie Carlson and Louis Moses (2001) demonstrated robust, statistically significant correlations between performance on false-belief tasks and classic measures of executive inhibition, such as the Bear/Dragon task, the Day/Night Stroop task, and the Dimensional Change Card Sort (DCCS). A child who cannot suppress sorting cards by shape when instructed to sort by color will almost certainly fail to suppress naming the pencils when asked about an unobservable past belief. This strong empirical linkage ignited a foundational debate: does the three-year-old child lack the conceptual architecture of belief itself, or do they possess the representational understanding while lacking the prefrontal inhibitory machinery required to suppress the overwhelming, seductive pull of physical reality?

6.2 Working Memory Load and Representational Updating

Beyond inhibitory control, the execution of the Smarties Task places significant structural demands upon working memory capacity and representational updating architectures. The child’s cognitive workspace must simultaneously maintain, update, and manipulate multiple high-dimensional informational tokens:

  • Token 1: The historical, baseline identity of the package (the conceptual category of Smarties sweets);
  • Token 2: The revealed physical contents (the empirical reality of the pencils);
  • Token 3: The temporal timeline of epistemic exposure (the sequence separating the past state of ignorance from the present state of knowledge);
  • Token 4: The agent-specific perspectives (who was present during the revelation versus who remained outside the room).

These disparate data points must not merely be retained; they must be dynamically bound together in the correct relational matrix. Working memory research reveals that three-year-old children have highly restricted processing capacities, typically limited to holding one or two active relational tokens simultaneously. When forced to bind Token 1 (past expectation) with Token 3 (temporal sequence) while inhibiting Token 2 (current reality), the fragile working memory system experiences systemic cognitive overload.

Experimental manipulations that artificially attenuate working memory demands—such as placing the container out of sight beneath a desk, thereby removing ongoing perceptual reinforcement—measurably reduce error rates in children on the cusp of competence. However, developmental studies indicate that even when memory demands are engineered to baseline minimums, children under 38 months continue to exhibit persistent failure rates. This confirms that while working memory constraints undoubtedly modulate real-time task performance, they operate in tandem with genuine conceptual limitations regarding the nature of mental representation.

6.3 Cognitive Flexibility and Perspective Shifting

The third pillar of the executive function suite modulating the Smarties Task is cognitive flexibility, characterized as the ability to disengage from one conceptual dimension and smoothly transition to another. In the unexpected contents paradigm, the child is compelled to execute rapid perspective shifts across two distinct theoretical axes: the diachronic (temporal) axis and the synchronic (interpersonal) axis.

On the diachronic axis, the child must shift back and forth between their own current cognitive state (knowing the reality) and their own past cognitive state (expecting candy). On the synchronic axis, the child must decouple from their immediate spatial-epistemic locus and adopt the perspective of a completely separate agent (the naive peer). This bidirectional shifting relies heavily upon the maturation of neural networks spanning the prefrontal cortex, specifically the dorsolateral prefrontal cortex (DLPFC), the ventromedial prefrontal cortex (vmPFC), and the anterior cingulate cortex (ACC).

Neurodevelopmental maturation across these prefrontal networks exhibits an accelerated period of synaptogenesis, axonal myelination, and dendritic pruning precisely between three and five years of age. Longitudinal investigations tracking executive function batteries and neuroimaging markers confirm that the structural and functional maturation of these prefrontal circuits directly underpins the emergence of the cognitive flexibility necessary to arbitrate between competing perspectives, establishing a bidirectional developmental highway wherein executive maturation scaffolds social mentalizing, and social engagement stimulates executive refinement.

7. Language, Pragmatics, and Communication in Task Performance

7.1 The Role of Syntactic Complement Clauses

One of the most theoretically profound contributions to understanding false-belief performance emerged from developmental linguistics, specifically through the groundbreaking work of Jill de Villiers (2000). De Villiers posited that the conceptual capacity to represent false beliefs does not develop independently of language, but is directly catalyzed and enabled by the acquisition of a specific, sophisticated grammatical structure: tensed syntactic complementation.

In linguistic syntax, complement clauses occur when a main clause containing a mental state or communication verb embeds a full, declarative proposition as its direct object. For example:

“The child thinks [that the box contains Smarties].”

The unique structural property of this grammatical architecture lies in its truth-conditional independence: the embedded complement proposition can be empirically false in the physical world without rendering the overall sentence false. If the box physically contains pencils, the proposition “the box contains Smarties” is flatly false. Yet, the matrix sentence—”The child thinks that the box contains Smarties”—remains completely, objectively true. De Villiers argued that this linguistic structure provides the exact representational scaffolding required to solve the Smarties Task. Before the child acquires the syntax of complementation, their cognitive system lacks the representational format necessary to encode an untruth without rejecting the entire cognitive proposition as an error.

Compelling empirical support for this linguistic scaffolding hypothesis stems from rigorous training studies. When young, transitional preschoolers are systematically trained on the grammar of sentential complements—even using non-mental communicative verbs such as “say” (e.g., “The boy said that the car is in the tree”)—their performance on subsequent unexpected contents tasks accelerates dramatically compared to control cohorts trained on standard complex syntax (such as relative clauses). Furthermore, cross-linguistic studies demonstrate that children acquiring languages that grammaticalize evidentiality—explicitly appending morphemes denoting whether information was learned via direct sight, hearsay, or inference—frequently exhibit accelerated false-belief competence, demonstrating the profound formative power of linguistic architecture on social cognition.

7.2 Conversational Pragmatics and Misleading Contexts

A crucial critique leveled against the standardized Smarties Task stems from the domain of conversational pragmatics. Led by cognitive scientist Michael Siegal, this perspective contends that children’s apparent epistemic failures do not reflect conceptual deficits, but rather an acute disconnect between the pragmatic conversational assumptions of the adult experimenter and the child participant.

Within human communication, as formalized by Paul Grice’s Cooperative Principle, interlocutors operate under mutual maxims of relevance, quantity, and quality. When an adult authority figure—an experimenter—holds up a familiar container, tricks a young child by revealing pencils, closes the container, and then asks: “What did you think was inside?”, the social context is fraught with pragmatic ambiguity. The child may perceive the adult’s query as fundamentally absurd if interpreted literally: why would an adult ask about a past, discarded error when the true, exciting reality of the pencils has just been unmasked? In the child’s estimation, naming the pencils is the most cooperative, informative, and reality-congruent answer to offer the adult.

To test this pragmatic interference hypothesis, researchers developed revised linguistic framings that aligned more transparently with children’s conversational schemas. When the pragmatic ambiguity is explicitly resolved—for instance, by asking, “When you first saw this box, before we opened it, did you think it had pencils or did you think it had Smarties?”—error rates decline significantly among transitional four-year-olds. However, while pragmatic refinements undeniably clean up peripheral performance noise, pure pragmatic theories cannot account for why young three-year-olds consistently fail the task across completely non-verbal, naturalistic paradigms, confirming the existence of a core representational threshold beneath pragmatic layers.

7.3 Semantic Development of Epistemic Verbs

The Smarties Task serves as an empirical crucible for observing the semantic crystallization of mental state vocabulary during the preschool years. Epistemic verbs such as know, think, guess, remember, and pretend do not emerge fully formed; they undergo an extended, turbulent semantic ontogeny.

Three-year-old children initially deploy cognitive verbs without representational precision, utilizing them primarily as conversational conversational devices, mood modulators, or expressions of certainty. For instance, a young child might say “I think so” merely to signal tentative agreement, or use “I know” simply to assert conversational dominance, rather than referencing distinct internal epistemic states. The critical semantic demarcation between knowing (which implies epistemic certainty tethered to true justified belief) and thinking (which permits subjective belief decoupled from truth) does not fully stabilize until around four years of age.

The administration of the unexpected contents paradigm provides an explicit operational index of this semantic stabilization. Numerous developmental studies reveal that the frequency and richness of maternal mental-state talk—parents who naturally discourse about mental states using nuanced epistemic verbs during shared book reading and everyday interactions—directly correlate with early success on the Smarties Task. When parents regularly articulate causal links between perceptual exposure and epistemic states (“We know it’s a dog because we saw it bark, but Daddy thinks it’s a cat because he didn’t look”), children internalize the semantic boundaries of epistemic terms years ahead of peers raised in linguistically impoverished environments.

8. Comparative Analysis: The Smarties Task versus The Sally-Anne Task

8.1 Structural Differences: Deceptive Containers versus Spatial Displacements

Although the Smarties Task and the Sally-Anne task are both classified as classic false-belief paradigms, their underlying cognitive affordances, stimulus structures, and processing demands diverge significantly. The Sally-Anne paradigm—derived from Wimmer and Perner’s original 1983 Maxi task—is fundamentally a dynamic spatial-displacement task. Its narrative requires tracking the movement of an external object across a spatial coordinate system, mediated by two distinct locations (the basket and the box) and contingent upon the spatial movement of agents entering and exiting a scene.

In sharp contrast, the Smarties Task is a static representational-conflict task. There is no spatial translocation; the target object remains anchored in a single, unmoving container throughout the entire protocol. The theoretical conflict does not unfold across physical space, but across conceptual space: the tension between the deceptive exterior semiotics of the container and its hidden internal physical contents. This static architecture provides decisive experimental advantages:

  • It completely eliminates narrative comprehension demands and story memory decay;
  • It removes spatial tracking, directional mapping, and location-memory burdens;
  • It relies upon immediate, real-world sensory interactions rather than the visual abstraction of puppet theatre or two-dimensional pictorial narratives.

Consequently, the Smarties Task strips away the extraneous mechanical baggage of the Sally-Anne design, isolating the child’s capacity to process the epistemic rupture between appearance and reality.

8.2 Testing Self-Belief: A Unique Paradigm Capability

The most consequential theoretical divergence between the two classic paradigms resides in their epistemic capacity: the Sally-Anne task is structurally incapable of assessing a child’s understanding of their own false beliefs. In Sally-Anne, the child participant sits outside the narrative as an omniscient observer. The child is never personally deceived; they watch the transfer unfold in real time. Therefore, the Sally-Anne paradigm can only evaluate third-person mentalizing.

The Smarties Task, by contrast, operates as a dual-vector diagnostic instrument, concurrently capturing both first-person (self) and third-person (other) false-belief attribution within a single, unified testing session. Because the child participant is actively misled by the commercial packaging, they experience an authentic, lived epistemic error. This enables developmentalists to investigate profound theoretical debates that cannot be approached via Sally-Anne:

  • Do children learn to understand their own minds before they understand other minds?
  • Does representational Theory of Mind emerge as a uniform, domain-general conceptual shift, or does it develop asymmetrically across self and other?

Empirical data generated by the Smarties paradigm revealed that children almost universally pass or fail both the self-belief and other-belief questions simultaneously. This extraordinary concordance provided critical empirical support for Theory-Theory (which claims that the child acquires a unified, abstract theoretical framework governing minds generally) over pure Simulation Theory (which posits that children primarily understand others by introspecting upon their own first-person experiences and projecting them outward).

8.3 Variability in Performance Profiles and Convergence Rates

In their comprehensive 2001 meta-analysis, Wellman, Cross, and Watson examined whether systemic performance discrepancies exist between location-change tasks (Sally-Anne) and unexpected-contents tasks (Smarties). Their statistical synthesis revealed a remarkable overall convergence: across hundreds of experimental variations, both paradigms capture the same underlying developmental watershed occurring between three and five years of age.

However, when fine-grained chronological increments are analyzed, subtle task-specific biases emerge. Some cohorts of children pass the Sally-Anne task slightly earlier than the Smarties Task, whereas others show the reverse pattern. The primary driver of this variation centers upon the presence of visual-perceptual salience. In the Smarties Task, once the pencils are placed back into the box, they are hidden from direct sight, which can marginally assist inhibitory control. Conversely, the packaging remains an active, visually overwhelming distractor that continually re-asserts its deceptive identity.

In the Sally-Anne task, the chocolate resides inside the real container, remaining hidden, but the physical presence of the two visible boxes can pull children toward pointing to the physical reality. Methodologists emphasize that because subtle variations in question phrasing, stimulus material, and inhibitory demands can induce transient performance splits, researchers evaluating Theory of Mind must never rely upon a single experimental design. Best practices dictate the deployment of a comprehensive, multi-task battery that combines unexpected-contents paradigms with unexpected-transfer paradigms to triangulate a child’s true underlying representational competence.

9. Methodological Variations and Experimental Manipulations

9.1 Deceptive Intent and the Trickery Manipulation

Recognizing the profound cognitive demands imposed by the standard Smarties Task, developmental researchers began experimenting with contextual modifications designed to investigate the psychological limits of children’s mentalizing. Among the most potent manipulations is the introduction of deceptive motivation, pioneered by Kate Sullivan and Ellen Winner (1993).

In this variation, the framing of the task is fundamentally altered from passive observation to active, conspiratorial mischief. Before introducing the naive peer, the experimenter invites the child to participate in an active prank: “Let’s play a trick on Johnny! Let’s take the Smarties out and hide these pencils inside, and see if we can fool him!” The child actively assists in packing the pencils into the confectionery tube and replacing the cap. When asked what Johnny will think is inside when he enters, children as young as three-and-a-half years old exhibit a dramatic surge in correct, false-belief answers.

The cognitive facilitation induced by this “trickery manipulation” is profound. By embedding the task within a competitive, playful social schema, several cognitive operations are immediately transformed:

  • The child’s attention is intentionally directed toward the discrepancy between appearance and reality;
  • The pragmatic ambiguity of the experimenter’s questioning is completely dispelled, as the explicit goal is to create an erroneous belief;
  • The social-evolutionary drive to execute deception provides motivational scaffolding that helps override the prepotent realist bias.

However, developmental boundaries persist: for children under 36 months, even the most vivid, conspiratorial trickery manipulation fails to rescue them from realist contamination. At that foundational age, children will happily participate in packing the pencils to “trick” their friend, but when asked what their friend will think, they still blurt out: “Pencils!”

9.2 Visual Aids, Physical Traces, and Explicit Anchoring

To systematically untangle whether three-year-olds fail the Smarties Task due to representational deficits or episodic memory decay of their initial mental state, Peter Mitchell and Hazel Lacohée (1991) devised an ingenious experimental modification known as the Posting-Box Paradigm. In this protocol, before the Smarties tube is opened, the experimenter asks the child to select a picture representing what they believe is inside the box from an array of images. The child selects a picture of Smarties and physically “posts” it into a secure, opaque slit in a wooden posting box.

The experimenter then reveals the atypical contents (pencils), confirms reality, and asks the critical self-belief question: “When you posted that picture into the box, what did you think was inside?” Mitchell and Lacohée discovered that the presence of this tangible, physical anchor drastically increased success rates among three-year-old children. By leaving an immutable physical trace of their prior epistemic state, the child was no longer forced to execute an abstract, unassisted retrieval of a vanished internal representation. The photograph served as an externalized cognitive scaffold that resisted the retroactive overwrite of current reality.

This discovery ignited intense theoretical debate. Did the posting-box task prove that three-year-olds actually possess a representational Theory of Mind that is simply masked by memory retrieval failure in the standard protocol? Or did the physical photograph simply transform the false-belief task into a simple, non-mentalistic visual memory test? Conceptual purists argued that the child was no longer reasoning about an invisible past belief, but merely reporting the physical contents of the posting box (“I posted the picture of Smarties”), highlighting the delicate empirical boundary between assessing meta-representational cognition and testing physical contingency recall.

9.3 Varying Deceptive Stimuli: Familiar Packaging versus Novel Objects

The universality and generalizability of the unexpected contents phenomenon have been extensively tested by systematically manipulating the physical properties of the deceptive stimuli. Across different global regions, researchers replaced the British Smarties tube with culturally relevant equivalents: iconic yellow Crayola crayon boxes containing small spoons or string; red Coca-Cola cans containing sand; or standard milk cartons containing small rocks.

These variations demonstrated that the strength of the initial expectation is directly modulated by the cultural semiotics of the package. Branded, highly commercial packaging—which carries rigid, deterministic category associations—generates stronger initial expectations, paradoxically intensifying the inhibitory control challenge during Phase 3 revelation. Conversely, generic or ambiguous containers (such as plain, unbranded tin cans) produce weaker baseline expectations, occasionally allowing children to equivocate during testing.

Furthermore, experimental psychologists expanded the unexpected contents paradigm beyond vision to encompass diverse sensory modalities. In tactile and haptic variations, children are handed an object that looks convincingly like a real, raw hen’s egg, only to discover through manual touch that it is carved from solid, heavy stone. In gustatory variations, children are shown what appears to be a cup of sweet orange juice, which is revealed to be solid orange gelatin. The relentless developmental consistency observed across all these sensory iterations confirms that the age-four shift is not an artifact of visual processing or packaging conventions, but represents a domain-general cognitive restructuring across all perceptual and representational systems.

10. Theoretical Debates: Representational Deficit versus Performance Limitations

10.1 The Representational Deficit Account (Theory-Theory)

The dominant theoretical paradigm emerging from the Smarties Task is the Theory-Theory framework, championed prominently by Josef Perner, Alison Gopnik, and Henry Wellman. Theory-Theory conceptualizes the developing child as an intuitive, informal scientist who navigates the world by constructing, testing, refining, and replacing internal theoretical frameworks of physical, biological, and psychological phenomena.

From the Theory-Theory perspective, the failure of three-year-olds on the unexpected contents task represents a genuine representational deficit. The young child is not simply suffering from transient attentional lapses, linguistic confusion, or poor memory. Rather, their internal “mentalizing theory” fundamentally lacks the conceptual construct of a “misrepresentation.” At three years of age, the child operates with a causal theory of human agency based on real-world situations (Situation Theory). They can conceptualize desires and direct perception because these constructs tie directly to physical reality: an agent desires an object, or an agent looks at an object.

Around age four, empirical anomalies accumulate to the point where the existing situation theory collapses under cognitive friction, triggering a paradigm shift analogous to scientific revolutions. The child constructs a new, higher-order theoretical framework: the Representational Theory of Mind. This new theory introduces an unobservable, mediating theoretical entity—the belief. With this conceptual innovation, the child understands that the mind does not copy reality; it constructs hypothetical models of reality that can be objectively false. Therefore, Theory-Theorists view the four-year-old transition as a radical, discontinuous conceptual leap in the child’s cognitive architecture.

10.2 Simulation Theory and Perspective-Taking

In direct opposition to the Theory-Theory framework, proponents of Simulation Theory—such as philosophers Alvin Goldman and Robert Gordon—argue that children do not understand human minds by deploying an abstract, scientific theory of folk psychology. Instead, Simulation Theory posits that humans understand other minds through direct, first-person imaginative empathy: we use our own internal cognitive and emotional apparatus as an offline, dynamic model to simulate the experiences of others.

When applied to the Smarties Task, Simulation Theory reinterprets children’s failures as a breakdown in imaginative quarantine. To predict what a naive peer will think when encountering the closed Smarties tube, the child must run a cognitive simulation. They must project themselves into the peer’s shoes, feed the peer’s perceptual inputs (seeing the closed box) into their own cognitive system, generate a simulated prediction, and read off the resulting belief. However, this simulation can only yield an accurate prediction if the child successfully quarantines—or suppresses—their own privileged knowledge that the tube actually contains pencils.

If the imaginative quarantine fails, the child’s own true knowledge leaks into the simulation run, contaminating the offline model. Under Simulation Theory, the three-year-old child fails the Smarties Task not because they lack the abstract concept of a belief, but because their imaginative simulation mechanisms are insufficiently isolated from their current reality. Over recent decades, many developmentalists have converged on hybrid models, acknowledging that children utilize both theoretical, rule-based deductions and affective, imaginative simulations depending upon the complexity and immediacy of the social context.

10.3 Nativism, Modularity, and The Two-Systems Theory

A third, highly influential theoretical perspective is grounded in cognitive nativism and modularity, spearheaded by Alan Leslie and, more recently, reformulated into Two-Systems Theories by Ian Apperly and Stephen Butterfill. Nativist modularity rejects the claim that a radical conceptual revolution occurs at age four. Instead, Leslie argues that the core conceptual machinery of Theory of Mind—the Theory of Mind Mechanism (ToMM)—is an innate neurodevelopmental module present from early infancy.

Why, then, do three-year-olds universally fail the explicit Smarties Task? Leslie attributes this failure to the immaturity of a non-modular, domain-general cognitive component termed the Selection Processor (SP). The Selection Processor is an inhibitory mechanism responsible for selecting the correct belief attribution from competing candidates. Because the physical pencils are perceptually real and hyper-salient, the Selection Processor of a three-year-old is simply too weak to suppress that reality, causing the child to fail the explicit test despite possessing an innate concept of belief.

Building upon this debate, Apperly and Butterfill introduced the Two-Systems account to reconcile the shocking discrepancies between infant implicit looking-time tasks and preschool verbal tasks:

  • System 1 (Implicit / Automatic): A phylogenetically ancient, fast, cognitively efficient, and encapsulated mentalizing system that operates automatically from early infancy. System 1 tracks gaze, simple actions, and visual perspectives, manifesting in anticipatory looking behaviors, but is rigid and incapable of processing complex, counterfactual metarepresentations.
  • System 2 (Explicit / Controlled): A phylogenetically modern, slow, cognitively demanding, and language-dependent mentalizing system that matures alongside the prefrontal cortex between three and five years of age. System 2 is flexible, handles complex syntactic complementation, enables explicit verbal reasoning about false beliefs, and powers success on classic tasks such as the Smarties Task.

Under this synthesis, the Smarties Task does not measure the dawn of social cognition from absolute zero; it evaluates the developmental consolidation of System 2, tracking the moment when implicit social tracking is translated into explicit, linguistically structured metacognitive awareness.

11. Cross-Cultural Replications and Environmental Variations

11.1 Cross-Cultural Universality of the Age-Four Shift

A vital imperative for developmental psychology is confirming whether empirical benchmarks established in Western, Educated, Industrialized, Rich, and Democratic (WEIRD) societies represent universal biological milestones or localized cultural artifacts. To resolve this question, an ambitious multinational collaborative study led by Tara Callaghan and colleagues (2005) administered standardized false-belief tasks—including the unexpected contents paradigm—across radically diverse cultural landscapes.

The researchers tested children across remote indigenous communities in Samoa, traditional hunting and gathering villages in Vanuatu, rural farming regions in Peru, rural communities in Thailand, and industrialized urban centers in Canada. The testing materials were meticulously adapted to ensure profound cultural familiarity: where children had never seen a cardboard Smarties tube, experimenters utilized locally recognized commercial containers or deceptive traditional containers (such as specialized hollow bamboo stalks concealing surprising insects or stones).

The empirical results revealed a profound, unequivocal cross-cultural invariance: despite dramatic disparities in formal pedagogical systems, literacy levels, communicative traditions, and socio-economic frameworks, children across every single culture demonstrated a transition from floor-level failure at age three to decisive, high-rate success around five years of age. While subtle local shifts emerged—such as children in certain traditional collectivist communities passing several months later than urban cohorts due to differential pedagogical styles—the developmental trajectory remained remarkably invariant. This universal cross-cultural shift provides powerful evidence that the emergence of a representational Theory of Mind is driven by a deep biological, neuro-maturational schedule shared by all human populations.

11.2 Socio-Environmental Modulators: Siblings and Family Dynamics

While the broader biological schedule of Theory of Mind is universally constrained, environmental factors exert measurable influence on the exact chronological velocity at which individual children master the Smarties Task. Among the most extensively documented socio-environmental influences is the Sibling Effect, first rigorously detailed by Josef Perner, Ted Ruffman, and Susan Leekam in their seminal 1994 paper, “Theory of mind is contagious: You can catch it from your brothers and sisters.”

Perner, Ruffman, and Leekam discovered that children growing up in households with older siblings pass the unexpected contents task significantly earlier—often by as much as six to eight months—than only children or first-born children. The mechanisms underpinning this sibling advantage are explicitly social-developmental:

  • Children with older siblings are constantly immersed in complex, competitive social conflicts, including teasing, trickery, playful deception, and negotiation;
  • Older siblings provide advanced models of communicative intent, frequently articulating counterfactual claims during imaginative and collaborative pretend play;
  • The younger child is routinely exposed to multi-party conversations where competing desires, perspectives, and false assumptions are overtly negotiated in real time.

In addition to sibling dynamics, parental discourse styles play a critical modulatory role. Mothers and fathers who utilize high levels of “mind-mindedness”—treating their young children as independent psychological agents possessing internal thoughts, and consistently talking about the causal links between mental states and behavior—raise children who master the Smarties Task significantly earlier than parental cohorts who communicate primarily through non-mentalistic behavioral directives.

11.3 Deaf Children of Hearing versus Deaf Parents

Perhaps the most poignant and scientifically illuminating natural experiment regarding environmental modulators of Theory of Mind involves studies of deaf children. Seminal research led by Candida Peterson and Michael Siegal (1995) examined performance on the Smarties and Sally-Anne tasks across two distinct cohorts of deaf children:

  • Deaf Children of Deaf Parents (Native Signers): Born into households where sign language is fluently spoken from birth, providing immediate, naturalistic communicative immersion;
  • Deaf Children of Hearing Parents (Late Signers): Comprising over 90% of deaf births, these children are born to parents who do not know sign language, resulting in severe early linguistic and communicative deprivation during the critical early preschool years.

The empirical divergence between these two groups was striking. Deaf native signers passed the deceptive box task at developmental rates identical to, and occasionally earlier than, hearing neurotypical children, effortlessly decoupling mental representation from reality by age four. In stark, tragic contrast, deaf late signers of hearing parents exhibited severe, persistent delays on the Smarties Task, frequently failing to grasp false beliefs until eight, ten, or even twelve years of age—mirroring the severe performance deficits observed in autistic cohorts.

Crucially, once these late-signing deaf children acquired fluent sign language and engaged in extended communicative discourse with peers, their Theory of Mind competence rapidly normalized. This double dissociation definitively isolated the paramount developmental truth: the emergence of a representational Theory of Mind is not dependent upon auditory processing or speech mechanics, but requires sustained, rich, early communicative immersion. Without access to everyday communicative discourse—the rich conversational flow where thoughts, misunderstandings, and varying perspectives are continually externalized—the internal cognitive construction of other minds is profoundly, severely compromised.

12. The Enduring Legacy and Contemporary Re-evaluations

12.1 Impact on Developmental Epistemology and Cognitive Science

The publication of Josef Perner, Susan Leekam, and Heinz Wimmer’s 1987 study firmly established the unexpected contents paradigm as one of the most transformative experimental breakthroughs in twentieth-century psychological science. By condensing the elusive, philosophical problem of other minds into a pristine, two-minute experimental protocol, they provided cognitive science with an indispensable diagnostic and theoretical tool. The paradigm demonstrated that cognitive development is not merely an incremental accumulation of facts or associative conditioning scripts; it is defined by qualitative conceptual revolutions that reconstruct how human beings perceive reality.

The cross-disciplinary ripple effects of the Smarties Task extend far beyond developmental psychology:

  • Philosophical Epistemology: The paradigm anchored philosophical debates surrounding folk psychology, intentionality, and mental realism in robust empirical methodology, demonstrating that the conceptual distinction between appearance and reality is an ontogenetic achievement rather than an innate philosophical given.
  • Evolutionary Anthropology: Comparative researchers deployed adapted deceptive-container assessments to non-human primates, corvids, and canines, mapping the phylogenetic evolutionary boundaries of representational mentalizing.
  • Artificial Intelligence and Robotics: As computational scientists strive to engineer autonomous agents capable of genuine social collaboration, the unexpected contents protocol serves as a foundational benchmark for programming cognitive architectures that can track dynamic, erroneous, and updating human beliefs.

12.2 Critiques from the Implicit Mentalizing Paradigm

Despite its foundational status, the classic Smarties paradigm has encountered formidable theoretical critiques over the past two decades, originating predominantly from the revolution in implicit mentalizing paradigms. In a historic 2005 study published in Science, Kristine Onishi and Renée Baillargeon deployed a novel violation-of-expectation (VoE) looking-time protocol, presenting evidence that 15-month-old infants looked significantly longer when an actor searched for a hidden toy in a location inconsistent with the actor’s false belief.

This explosive finding—followed by a proliferation of anticipatory-looking and pupil-dilation studies—ignited fierce debates that continue to polarize developmental science. Nativists argued that if 15-month-old infants already grasp false beliefs, then explicit verbal tasks such as the Smarties Task are fundamentally flawed, merely measuring extraneous performance demands such as speech, executive inhibition, and working memory, rather than genuine mentalizing competence.

However, the implicit mentalizing revolution has recently encountered severe empirical pushback during the wider behavioral replication crisis. Multiple high-powered, multi-laboratory pre-registered replication attempts have systematically failed to replicate prominent infant false-belief looking-time effects. Furthermore, prominent cognitive theorists contend that infant looking behaviors can be fully explained by low-level perceptual heuristics, such as novelty detection, perceptual affinity, and spatial tracking mechanisms, completely devoid of representational attribution. Consequently, the explicit Smarties Task has retained its enduring, robust validity. It remains the gold standard because it demands that the child explicitly, unambiguously decouple mental representation from ontological truth, providing incontrovertible proof of representational agency.

12.3 Future Trajectories in Neuroimaging and Artificial Social Cognition

As cognitive science marches into the twenty-first century, the conceptual architecture engineered by Perner, Leekam, and Wimmer continues to inspire cutting-edge empirical trajectories across neuroimaging and artificial intelligence. In modern developmental cognitive neuroscience, researchers utilize high-density functional Near-Infrared Spectroscopy (fNIRS) and child-friendly functional Magnetic Resonance Imaging (fMRI) to map the real-time neural correlates of the Smarties Task in early childhood.

These neuroimaging studies have pinpointed a specialized, highly integrated brain network known as the Mentalizing Network, encompassing:

  • The Right and Left Temporoparietal Junction (TPJ), heavily implicated in the transient computation of distinct, temporary mental states and the execution of epistemic decoupling;
  • The Medial Prefrontal Cortex (mPFC), engaged in abstract social reasoning, perspective evaluation, and attributing stable personality and cognitive traits to social agents;
  • The Precuneus and Posterior Cingulate Cortex, responsible for autobiographical memory retrieval and contextual perspective synthesis.

Neuroimaging tracks how functional connectivity across these disparate nodes undergoes dramatic structural consolidation precisely between three and five years of age, physically mirroring the behavioral transition observed in the Smarties Task.

Simultaneously, the Smarties paradigm has emerged as a critical evaluation benchmark in contemporary Artificial Intelligence. As Large Language Models (LLMs) and autonomous neural architectures achieve breathtaking linguistic fluency, computer scientists subject these models to complex, textually modified unexpected contents assessments. While modern LLMs can often predict the correct verbal string (“Smarties”) due to vast statistical pattern matching over trillions of parameters, cognitive scientists debate whether this performance reflects authentic, mechanistic representational understanding or merely hyper-sophisticated, non-mentalistic syntactic regurgitation. The conceptual rigor of the unexpected contents protocol remains the primary empirical boundary stone separating true human cognitive architecture from synthetic behavioral mimicry.

Conclusion

The deceptive box paradigm conceived by Josef Perner, Susan Leekam, and Heinz Wimmer in 1987 represents one of the most brilliant and enduring methodologies in the history of cognitive science. By transforming the abstract, intractable philosophical dilemma of epistemic representation into a simple, elegant interaction with a confectionery tube and five pencils, their experiment permanently illuminated the mechanics of the human mind. The Smarties Task unmasked the remarkable ontogenetic revolution through which the human child ceases to inhabit an unmediated reality, learning to navigate the luminous, complex, and subjective landscape of mental representations. In doing so, Perner, Leekam, and Wimmer provided humanity with an indelible cognitive mirror, proving that the capacity to understand our own errors, to empathize with the misinformed, and to comprehend the diverse, invisible perspectives of our fellow beings is the foundational cornerstone of the human social mind.

References

  • Baron-Cohen, S., Leslie, A. M., & Frith, U. (1985). Does the autistic child have a “theory of mind”? Cognition, 21(1), 37–46. https://doi.org/10.1016/0010-0277(85)90022-8
  • Bennett, J. (1978). Some remarks about concepts. Behavioral and Brain Sciences, 1(4), 557–560. https://doi.org/10.1017/S0140525X0007666X
  • Callaghan, T., Rochat, P., Lillard, A., Shapiro, D., Guo, J., Cohen, S., Run, C., & Naka, M. (2005). Synchrony in the onset of mental-state reasoning: Evidence from five cultures. Psychological Science, 16(5), 378–384. https://doi.org/10.1111/j.0956-7976.2005.01544.x
  • Carlson, S. M., & Moses, L. J. (2001). Individual differences in inhibitory control and children’s theory of mind. Child Development, 72(4), 1032–1053. https://doi.org/10.1111/1467-8624.00333
  • de Villiers, J. G. (2000). Language and theory of mind: What are the developmental relationships? In S. Baron-Cohen, H. Tager-Flusberg, & D. J. Cohen (Eds.), Understanding other minds: Perspectives from developmental cognitive neuroscience (2nd ed., pp. 83–123). Oxford University Press.
  • Dennett, D. C. (1978). Beliefs about beliefs. Behavioral and Brain Sciences, 1(4), 568–570. https://doi.org/10.1017/S0140525X00076701
  • Gopnik, A., & Astington, J. W. (1988). Children’s understanding of representational change and its relation to the understanding of false belief and the appearance-reality distinction. Child Development, 59(1), 26–37. https://doi.org/10.2307/1130386
  • Leslie, A. M. (1987). Pretense and representation: The origins of “theory of mind”. Psychological Review, 94(4), 412–426. https://doi.org/10.1037/0033-295X.94.4.412
  • Mitchell, P., & Lacohée, H. (1991). Children’s early understanding of false belief. Cognition, 39(1), 17–27. https://doi.org/10.1016/0010-0277(91)90033-B
  • Onishi, K. H., & Baillargeon, R. (2005). Do 15-month-old infants understand false beliefs? Science, 308(5719), 255–258. https://doi.org/10.1126/science.1107621
  • Perner, J. (1991). Understanding the representational mind. The MIT Press.
  • Perner, J., Leekam, S. R., & Wimmer, H. (1987). Lack of belief about belief: Young children’s and autistic people’s failure to understand both the true and false belief of other people. Cognition, 26(3), 283–309. https://doi.org/10.1016/0010-0277(87)90011-4
  • Perner, J., Ruffman, T., & Leekam, S. R. (1994). Theory of mind is contagious: You can catch it from your brothers and sisters. Child Development, 65(4), 1228–1238. https://doi.org/10.2307/1131316
  • Peterson, C. C., & Siegal, M. (1995). Deafness, conversation and theory of mind. Journal of Child Psychology and Psychiatry, 36(3), 459–474. https://doi.org/10.1111/j.1469-7610.1995.tb01303.x
  • Premack, D., & Woodruff, G. (1978). Does the chimpanzee have a theory of mind? Behavioral and Brain Sciences, 1(4), 515–526. https://doi.org/10.1017/S0140525X00076518
  • Siegal, M., & Beattie, K. (1991). Where to look for children’s knowledge: The underestimation of cognitive performance on false belief tasks. Cognition, 40(3), 247–264. https://doi.org/10.1016/0010-0277(91)90001-K
  • Sullivan, K., & Winner, E. (1993). Three-year-olds’ understanding of a false belief: The trick of deception. Cognition, 47(2), 95–111. https://doi.org/10.1016/0010-0277(93)90027-C
  • Wellman, H. M., Cross, D., & Watson, J. (2001). Meta-analysis of theory-of-mind development: The truth about false belief. Child Development, 72(3), 655–684. https://doi.org/10.1111/1467-8624.00304
  • Wimmer, H., & Perner, J. (1983). Beliefs about beliefs: Representation and constraining function of wrong beliefs in young children’s understanding of deception. Cognition, 13(1), 103–128. https://doi.org/10.1016/0010-0277(83)90004-5

Rate This Content

0.0 / 5 0 votes

Cite This Article

memjavad (2026, September 12). The Smarties Task (False Belief) – Josef Perner, Susan Leekam, and Heinz Wimmer. PSYCHOLOGICAL DATABASE. https://en.arabpsychology.com/experiments/smarties-task-false-belief-perner-leekam-wimmer/
memjavad. “The Smarties Task (False Belief) – Josef Perner, Susan Leekam, and Heinz Wimmer.” PSYCHOLOGICAL DATABASE, 12 September 2026, https://en.arabpsychology.com/experiments/smarties-task-false-belief-perner-leekam-wimmer/.
memjavad. “The Smarties Task (False Belief) – Josef Perner, Susan Leekam, and Heinz Wimmer.” PSYCHOLOGICAL DATABASE. September 12, 2026. https://en.arabpsychology.com/experiments/smarties-task-false-belief-perner-leekam-wimmer/.