For centuries, the nature of human cognition and the origins of our understanding of the external universe have stood as central battlegrounds in philosophy, epistemology, and developmental psychology. In his classic formulation, William James famously characterized the perceptual world of the human neonate as a “blooming, buzzing confusion”—an anarchic sensory kaleidoscope void of structural coherence, ontological stability, or causal order. Under this classic view, and within the later constructivist paradigm established by Jean Piaget, the infant was presumed to enter existence as an uncalibrated sensory tabula rasa, incapable of conceptualizing objects outside the immediate purview of sensory impressions and reflex motor actions. For the first half of the twentieth century, human intelligence was understood as an excruciatingly slow, emergent synthesis achieved entirely through gradual sensorimotor exploration and inductive associative learning.
This long-standing empirical orthodoxy was decisively upended in the late twentieth century by the pioneering research program of Elizabeth Spelke and her collaborators. Through a methodologically revolutionary suite of looking-time paradigms and theoretical innovations, Spelke demonstrated that human infants do not enter the world conceptually blind, nor do they inhabit a fragmented perceptual wilderness. Instead, preverbal infants possess an early-emerging, phylogenetically ancient, and domain-specific set of cognitive mechanisms known as “Core Knowledge.” Among these foundational domains, the core knowledge of physics operates as an innate inductive engine, enabling infants to carve the dynamic physical world into bounded, enduring entities that obey rigid mechanical constraints long before they can utter their first word, grasp an object with fine-motor precision, or crawl across a room.
Spelke’s framework posits that human infants are endowed with an evolutionary inheritance of foundational physical axioms: principles of cohesion, continuity, solidity, and contact. These principles do not operate as transient perceptual filters, nor do they mirror adult formal scientific mechanics; rather, they serve as non-verbal inferential systems that parse visual flux into individual three-dimensional objects, calculate continuous trajectories across space and time, and predict mechanical collisions. This article provides an exhaustive, multidisciplinary analysis of Elizabeth Spelke’s theory of core knowledge in infant physics. Across twelve comprehensive dimensions, we will trace the epistemological roots of nativism, dissect the empirical mechanics of infant physical perception, evaluate historical and methodological controversies, examine comparative evolutionary substrates, and investigate the profound implications of Spelkean intuitive physics for modern developmental neuroscience and artificial intelligence.
1. Foundations of Core Knowledge Theory and Elizabeth Spelke’s Epistemological Framework
1.1 Historical Emergence: Nativism Versus Radical Empiricism
The philosophical pedigree of core knowledge theory traces directly back to the classic debates between rationalist nativism and radical empiricism. In the seventeenth and eighteenth centuries, John Locke and David Hume championed an epistemic framework wherein the human mind begins as a clean slate (tabula rasa). For Locke, all ideas, from the simplest physical impressions to the most complex abstract relations, derived strictly from sensory experience and subsequent reflection. Hume extended this skepticism to mechanical causality itself, arguing that human observers never directly perceive physical necessity or causal force; rather, they observe spatiotemporal contiguity and constant conjunction, generating a psychological habit of expectation through repeated exposure. In direct contrast, René Descartes and Immanuel Kant asserted that sensory impressions are fundamentally uninterpretable in the absence of antecedent, innate organizing principles. Kant argued that concepts such as space, time, and causality do not emanate from sensory data; rather, they represent a priori forms of intuition and categories of the understanding that make sensory experience intelligible in the first place.
Throughout the early and mid-twentieth century, mainstream psychology largely operationalized radical empiricism through behaviorism and Piagetian genetic epistemology. Behaviorist models formulated by B.F. Skinner and John B. Watson rejected internal representational states altogether, attempting to explain cognitive development through stimulus-response associations, passive conditioning, and physical reinforcement schedules. Concurrently, Jean Piaget offered a more sophisticated, constructivist alternative: while acknowledging the active role of the child, Piaget maintained that neonates possess nothing more than basic sensorimotor reflexes (such as sucking and grasping). For Piaget, physical concepts—including the fundamental permanence, solidity, and spatial independence of physical bodies—were laboriously constructed over months of physical action, perceptual coordination, and the dual mechanisms of assimilation and accommodation. The infant mind was deemed devoid of internal representational architecture during early ontogeny.
Elizabeth Spelke revolutionized this theoretical landscape by re-evaluating infant cognitive architecture through the lens of modern evolutionary developmental biology and Chomskyan generative linguistics. Drawing an explicit epistemological parallel to Noam Chomsky’s universal grammar—which posits that children possess an innate language acquisition device that constrains and guides grammatical induction—Spelke argued that human infants inherit domain-specific cognitive mechanisms designed to solve universal ecological challenges. In the physical domain, natural selection could not afford to leave the fundamental parsing of the physical environment to the slow, error-prone vagaries of trial-and-error associative learning. An organism incapable of parsing cohesive entities, tracking moving predators, or understanding that a falling cliff or collapsing surface poses an existential threat would suffer catastrophic fitness costs. Thus, Spelke revived Kantian rationalism within an evolutionary framework, establishing that the initial state of human cognition is populated by structured, domain-specific inductive biases regarding the physical universe.
1.2 The Initial State Hypothesis and Domain-Specific Cognitive Architecture
At the center of Spelke’s theoretical architecture lies the “Initial State Hypothesis.” This hypothesis posits that human infants possess an innate foundational core of physical concepts that is present at the very beginning of postnatal life, prior to significant motor experience or physical manipulation. These core representations do not represent immature approximations of adult sensory perceptions, nor do they mirror the domain-general, associative learning networks popularized by connectionist architectures. Instead, core knowledge systems are domain-specific, modular neurocognitive mechanisms dedicated to specific ontological categories, such as inanimate physical objects, intentional animate agents, numerical magnitudes, and geometric spatial layouts.
To qualify as a genuine system of core knowledge within Spelke’s taxonomy, a cognitive mechanism must satisfy several rigorous diagnostic criteria. First, it must demonstrate early emergence; the foundational principles must be detectable as early in ontogeny as the limits of behavioral testing permit, typically within the first few weeks or months of life, before infants have acquired independent physical locomotion or manual dexterity. Second, a core knowledge system must exhibit phylogenetic continuity; because these systems represent ancient evolutionary adaptations designed to navigate the physical earth, homologous mechanisms should be detectable in non-human animals, such as non-human primates and birds, operating under similar ecological constraints. Third, core systems are characterized by task independence; the underlying principles operate automatically across an expansive array of contexts, unaffected by changes in explicit conscious task demands.
Crucially, core knowledge systems diverge sharply from classic Fodorian modularity in several fundamental dimensions. While Jerry Fodor envisioned mental modules as rigid, encapsulated input analyzers tied directly to specific sensory transducers, Spelke conceives of core physical systems as conceptual inference engines that operate across multiple perceptual modalities. Core knowledge of physics takes inputs from vision, audition, and haptics, translating sensory inputs into abstract, language-like mechanical representations. Furthermore, unlike domain-general cognitive architectures—which presume that a unified learning mechanism governs every domain from spatial mapping to syntactic parsing—Spelke’s domain-specific systems function autonomously. The inferential rules applied by the infant to predict the trajectory of a rolling stone do not apply to the behavior of a human caregiver or the evaluation of musical cadence, demonstrating a clean ontological fracture carved into the human mind by natural selection.
1.3 Methodological Revolutions in Investigating Preverbal Cognition
The primary barrier to uncovering the infant’s conceptual understanding of physics had historically been methodological. Piaget’s empirical reliance on manual search tasks—such as requiring an infant to lift a cloth barrier to retrieve a hidden toy—conflated cognitive conceptual competence with motor execution capacities. Because the human neuromuscular system develops along a strict cephalocaudal and proximodistal trajectory, young infants do not possess the requisite motor planning, bimanual coordination, or frontal inhibitory control necessary to reach for, grasp, and unveil concealed objects until approximately seven to nine months of age. Piaget mistakenly interpreted this physical reaching failure as an ontological void, concluding that hidden objects ceased to exist for the infant.
To penetrate this motor bottleneck, Spelke, alongside contemporaries such as Robert Fantz and Renée Baillargeon, pioneered the use of looking-time paradigms, specifically the Violation-of-Expectation (VoE) technique and habituation-dishabituation protocols. The fundamental theoretical premise underlying the VoE paradigm rests upon an established evolutionary and cognitive phenomenon: visual attention is intrinsically linked to novelty and epistemic surprise. When an organism encounters a visual display that accords with its internal model of physical reality, attention wanes rapidly through habituation. Conversely, when an environmental event violates foundational mechanical principles—such as an object passing through a solid wall or disappearing into absolute nothingness—the internal cognitive model is violated, triggering heightened attentional orienting, prolonged foveal fixation, and increased pupillary dilation.
The operationalization of looking-time paradigms required exquisite methodological controls to decouple conceptual mechanical violations from low-level perceptual novelties. In a prototypical Spelkean experiment, infants are familiarized with an apparatus or habituated to a standard kinematic event. Subsequently, they are presented with two test events: an “expected” (physically possible) event and an “unexpected” (physically impossible) event. Methodologically, the impossible event is intentionally designed to be perceptually more similar to the habituation event, or visually less complex, than the possible event. If infants fixate significantly longer on the impossible event despite its reduced superficial novelty, researchers can infer that the heightened visual allocation is driven not by low-level sensory salience, but by an internal cognitive representation of a mechanical impossibility. Through this ingenious methodological inversion, looking time was transformed into an objective epistemic readout of the preverbal human mind.
2. The Ontological Status of Physical Objects in Early Infancy
2.1 Defining the Infant’s Conception of Objecthood
What constitutes an “object” in the cognitive ecosystem of a three-month-old human infant? In common vernacular, objects are often characterized by their functional utilities, cultural meanings, or linguistic labels. For Elizabeth Spelke, however, the infant’s initial definition of an object is strictly ontological and mechanical. Through extensive empirical mapping, Spelke established that core physics defines an object as a bounded, coherent, three-dimensional physical entity that maintains its internal structural integrity and occupies a continuous region of space and time. This foundational definition establishes a profound cognitive bifurcation between genuine physical objects and other entities within the visual field.
Infants do not treat all physical phenomena identically. A critical accomplishment of Spelke’s experimental program was demonstrating that young infants systematically differentiate between unitary physical objects, non-cohesive substances, spatial collections, and background environmental surfaces. For example, when infants are presented with an aggregate mass of loose sand, pouring liquid, or a pile of disparate wooden blocks, their visual tracking and individuation systems do not deploy the same mechanical constraints that are automatically triggered by a discrete, solid wooden cube. Liquids and granular media deform, disperse, and pass through barriers without eliciting the looking-time spikes indicative of core violation. This demonstrates that core physics is not a generic theory of matter, but a highly targeted computational system explicitly tuned to the mechanics of macroscopic, bounded bodies.
Spelke’s criteria for individuating objects within complex visual arrays rely upon explicit spatial and kinematic boundaries. An entity is parsed as a distinct physical object if its component parts move in strict synchrony relative to one another, while exhibiting spatial discontinuity from adjacent surfaces. In visual scenes where multiple physical objects overlap, lean against one another, or share continuous boundaries, the core cognitive system does not rely on static semantic classifications. Instead, it utilizes structural geometry: surfaces that are internally connected, share continuous depth planes, and exhibit uniform displacement are integrated into a single object representation. Conversely, regions characterized by spatial gaps, depth discontinuities, or independent vectors of movement are immediately segmented into separate physical tokens.
2.2 Spatiotemporal Priority over Surface Features
One of the most profound discoveries emerging from Spelke’s laboratory is the absolute primacy of spatiotemporal information over surface featural properties during early ontogenetic object processing. When adults encounter physical objects, they instantly identify, categorize, and track them using a dense tapestry of featural dimensions, including color, chromatic saturation, texture, visual pattern, and fine geometric curvature. An adult seeing a red sphere and a green cube can effortlessly track their distinct trajectories, using color and shape as unambiguous diagnostic tags of identity.
Astonishingly, empirical investigations conducted by Spelke, Susan Carey, and Fei Xu revealed that infants under ten to twelve months of age exhibit striking limitations in utilizing surface features to individuate physical objects. In classic experimental demonstrations, an infant observes a screen. A blue, textured, cylindrical object emerges from the left side of the screen, moves across the stage, and retreats behind the screen. Subsequently, a bright yellow, smooth, triangular prism emerges from the right side and returns. To an adult observer, this display unequivocally confirms the existence of two numerically distinct objects. However, when the screen is removed to reveal either one or two objects, infants under ten months of age show no surprise upon seeing only a single object—provided the temporal cadence and spatial trajectory could be accounted for by a single entity executing a round-trip path.
Infants only infer the existence of two distinct objects when spatiotemporal boundaries are explicitly revealed—such as observing both objects simultaneously on opposite sides of the screen, or seeing one object disappear while another appears with a temporal gap that makes a continuous single trajectory physically impossible. This spatiotemporal priority aligns with the functional segregation of visual processing in the human brain. The dorsal visual pathway (“where/how” stream), which processes spatial relationships, volumetric boundaries, and kinematic trajectories, matures substantially faster than the ventral visual pathway (“what” stream), which resolves fine-grained chromatic and textural identification. Core physics is fundamentally anchored in the dorsal stream’s spatial-kinematic computations. Featural identification is a late-developing layer that only gradually interfaces with the infant’s primary physical token system.
2.3 The Indexing Model and the Physical Token System
To mathematically and mechanistically formalize how preverbal infants track objects through dynamic scenes, cognitive scientists have synthesized Spelke’s empirical findings with Zenon Pylyshyn’s Visual Indexing Theory (FINST, or “Fingers of Instantiation”). According to the indexing model, the infant visual architecture deploys a discrete set of mental pointers or “object tokens” that latch directly onto spatial entities in the visual field. These tokens do not initially contain detailed descriptive predicates such as “is red,” “is wooden,” or “is spherical.” Rather, an object token functions as a pure, referential, deictic pointer that tracks “Object-1,” “Object-2,” or “Object-3” as they traverse dynamic spatiotemporal coordinates.
This physical token system operates under strict structural capacity limits. Groundbreaking studies tracking infant looking time and manual retrieval limits demonstrate that the core object indexing system in infants possesses a strict capacity ceiling of approximately three to four concurrent items. If an infant observes one, two, or three items hidden inside an opaque container, they will systematically search until the exact number of objects is retrieved. However, if the array exceeds four items (for instance, presenting an infant with four versus eight items in an unsegmented presentation), the continuous tracking system suffers catastrophic degradation. The infant frequently fails to search at all, indicating that the core indexing system does not function as an analog magnitude estimator, but as a discrete, tokenized tracking engine that fails completely when its working memory slots are overloaded.
The existence of this tokenized indexing system explains how core physical representations support rapid, online kinematic predictions. When an object token is instantiated, the core knowledge system instantly applies the mechanical axioms of physics to that specific token. If Object-Token-1 moves behind an occluder, the system calculates an expected point of reappearance based on its velocity and vector, anticipating its emergence on the opposite side. If Object-Token-1 collides with Object-Token-2, the system projects the immediate deceleration of the impactor and the acceleration of the target. These inferences are computed instantaneously, enabling infants to allocate anticipatory gaze precisely to where an object should reappear or land, transforming visual tracking from a passive sensory reaction into an active predictive simulation of physical reality.
3. The Principle of Cohesion: Mechanics of Connected and Bounded Entities
3.1 Theoretical Formulation of the Cohesion Principle
The first axiomatic pillar in Elizabeth Spelke’s mechanics of the infant mind is the Principle of Cohesion. Formally defined, the cohesion principle dictates that objects are internally connected, bounded physical bodies that maintain their structural integrity as they move through space and time. Under this cognitive constraint, a physical object cannot spontaneously rupture into disparate, flying fragments, nor can separate, unconnected spatial entities spontaneously fuse into a single unitary body without the application of external destructive or binding forces. Cohesion ensures that an object’s surfaces retain their spatial connectivity during displacement.
The operational logic of the cohesion principle establishes an explicit prohibition within infant intuitive physics: spontaneous fission and fusion are treated as mechanical impossibilities. In the natural macroscopic world, a rock does not spontaneously transform into two distinct pebbles mid-flight; a bird does not dissolve into independent anatomical segments as it glides between trees. The infant core knowledge system codifies this environmental invariance into an inflexible inferential rule. An object must move as a unified whole; all points on the surface of the object must maintain their relative topological proximity unless an external physical interaction interrupts this structural coherence.
Furthermore, the cohesion principle marks the conceptual dividing line between rigid or semi-rigid solids and deformable matter or particulate aggregates. Because cohesion requires that an entity retains its boundaries during transit, infants anticipate that if any segment of an object is moved, the remainder of the object will move along the same vector. Conversely, when infants encounter substances that lack internal cohesive bonds—such as water, viscous fluids, or loose piles of sand—they do not deploy cohesion-based expectations. Core physics relies on the assumption of structural integrity; it treats cohesive objects as the ontological baseline of physical reality, relegating non-cohesive substances to a distinct ontological class that requires alternative mechanical logic.
3.2 Empirical Evidence from Violation-of-Expectation Experiments
The empirical verification of the cohesion principle was systematically executed through a landmark series of violation-of-expectation experiments designed by Elizabeth Spelke and her research team. In a quintessential experimental design, infants as young as three to four months of age were presented with a visually bounded entity resting upon a stage. An experimenter’s hand entered the display, grasped the top portion of the entity, and pulled it upward. In the cohesive (possible) condition, the entire object moved upward as an integrated, unitary mass. In the non-cohesive (impossible) condition, the object miraculously split in half: the top portion elevated with the hand, while the bottom portion remained stationary on the stage floor, revealing a clean, unannounced internal rupture.
The experimental results were unambiguous. Across diverse trials and counterbalanced sequences, infants looked significantly longer at the event in which the object spontaneously cleaved into pieces compared to the event in which the object moved as a cohesive whole. To confirm that the infants’ heightened looking time was not an artifact of seeing two separate shapes instead of one (a purely perceptual novelty), Spelke introduced baseline controls where the entity was already visibly divided into two separate, stationary blocks before the lift occurred. When infants were aware that two distinct blocks were present from the outset, the upward movement of only the top block elicited no surprise. The violation-of-expectation occurred exclusively when a presumed single, cohesive entity suffered spontaneous structural disintegration.
Subsequent investigations deepened these insights by testing cohesion violations under complex occluded conditions versus fully visible trajectories. When an object traversed behind an occluding screen and emerged as two independent fragments moving along divergent vectors, infants exhibited prolonged visual fixation upon the screen’s removal, demanding to inspect the internal state of the object. Furthermore, when researchers presented infants with piles of non-cohesive matter—such as sand or piles of unattached beads—infants did not express surprise when an implement picked up a portion of the pile while leaving the rest behind. Their looking-time behavior revealed an acute sensitivity to the boundaries of the cohesion constraint: it applies rigidly to bounded, discrete physical objects, but is appropriately suppressed when evaluating amorphous, aggregate matter.
3.3 Parsing Ambiguous Visual Scenes Through Cohesive Motion
One of the most complex visual challenges facing any biological or artificial vision system is the problem of scene segmentation: how does an observer determine where one object ends and another begins when surfaces overlap, cast shadows, and obscure one another? In 1983, Philip Kellman and Elizabeth Spelke executed what has become one of the most famous experiments in cognitive science: the classic “rod-and-block” demonstration. This experiment definitively established that preverbal infants utilize cohesive, common motion as the fundamental spatial primitive to parse ambiguous visual arrays.
Four-month-old infants were habituated to a visual display featuring a continuous, textured rectangular block. Directly behind this block rested a long, angled rod, such that only the top and bottom ends of the rod were visible, protruding from opposite sides of the central occluding block. Crucially, the rod moved back and forth horizontally behind the stationary block, moving in strict common motion (synchronous velocity, direction, and phase). The visual input itself was ambiguous: the visible scene contained two separate visual patches (the top rod-tip and the bottom rod-tip) separated by a solid barrier. Did the infants perceive two disconnected rod fragments moving in tandem, or did they perceive a single, continuous, unified rod whose center was temporarily occluded?
Following habituation, the central occluding block was removed entirely, and infants were presented with two alternating test events: a single, intact, continuous rod, or two broken rod segments separated by an empty gap where the block had previously been. If infants relied purely on low-level, two-dimensional sensory features—retinal patterns of visible color and line endings—the broken segments should have been familiar, and the continuous rod should have appeared novel. In striking contrast, the infants looked significantly longer at the broken rod. They had inferred that the two visible extremities belonged to a single, continuous, cohesive object unified behind the occluder, solely on the basis of their shared, synchronous motion. When the same experiment was performed with a stationary rod and block display, infants showed no preference, demonstrating that static visual alignment alone is insufficient for early object unit formation. It is the dynamic mechanics of common motion—the operationalization of cohesion—that allows the infant mind to solve the riddle of visual occlusion.
4. The Principle of Continuity: Spatiotemporal Paths and Existence
4.1 Formulation of the Continuity Constraint
The second foundational axiom of core physics is the Principle of Continuity. This constraint dictates that objects move only on connected, unbroken paths through space and time. Under the continuity principle, a physical object cannot simply vanish from one spatial coordinate and instantaneously materialize at another coordinate without traversing all intermediate points along the trajectory. Spatiotemporal continuity forms the bedrock of physical ontology; it guarantees that matter cannot undergo teleportation, that physical entities possess enduring identity through displacement, and that an object continues to exist even when it is temporarily absent from the visual field.
A critical corollary of the continuity principle is the infant’s distinction between perceptual absence and ontological cessation. When an object passes behind an opaque wall, a blanket, or a visual barrier, the optical rays reflecting from the object’s surface are extinguished at the observer’s retina. For a purely sensory, behaviorist, or radical empiricist architecture, the object ceases to exist the exact moment visual sensations cease. Under Spelke’s continuity constraint, however, the infant cognitive system executes a profound conceptual calculation: the termination of optical input is interpreted as an instance of spatial occlusion, not physical annihilation. The infant represents the unperceived entity as continuing to occupy a precise, albeit invisible, spatial coordinate.
The continuity constraint enforces a strict spatiotemporal calculus upon the infant visual engine. If an object is tracked entering the left side of an occluder, the system calculates a continuous, hidden trajectory. The object is expected to remain behind the occluder for a duration proportional to its velocity and the barrier’s width, and it must emerge predictably from the opposite edge. Any disruption in this continuous path—such as an instantaneous jump across a spatial divide, a failure to appear in an unoccluded spatial gap, or an emergence that occurs too early or too late—constitutes an explicit violation of continuity, firing an immediate cognitive alarm within the infant’s looking behavior.
4.2 Dismantling Piaget’s Concept of Object Permanence
For more than four decades, developmental psychology was dominated by Jean Piaget’s constructivist thesis that infants do not achieve a cognitive representation of object permanence until approximately eight to twelve months of age (during the transition through Stage IV of the sensorimotor period). Piaget based this profound conclusion upon the infamous “manual search deficit.” When a four- or five-month-old infant is reaching for an attractive toy, and the experimenter places an opaque cloth over the toy, the infant immediately halts their reaching action, often looking away or showing distress. To Piaget, this manual cessation proved that out of sight was literally out of mind: the infant lacked internal mental representations capable of sustaining the object’s existence across time and space.
The empirical research of Elizabeth Spelke, working alongside Renée Baillargeon, fundamentally dismantled this Piagetian dogma. Recognizing that manual searching imposes crushing demands on infant motor planning, motor inhibition, and means-end coordination, these researchers deployed looking-time and violation-of-expectation paradigms to assess whether infants represent hidden objects when the motoric requirement is eliminated. The findings reshaped the discipline: infants as young as 2.5 to 3.5 months of age demonstrate robust, undeniable representations of object permanence and spatiotemporal persistence.
In Baillargeon’s classic “drawbridge” (rotating screen) experiment, infants were habituated to an opaque screen that rotated back and forth through a 180-degree arc flat against a table, like a drawbridge. A solid wooden box was then placed visibly behind the path of the screen. In the possible event, the screen rotated upward, occluded the box, and halted precisely at the point where it would physically make contact with the solid box (e.g., stopping at 112 degrees). In the impossible event, the screen rotated upward, occluded the box, but continued its full 180-degree trajectory, laying completely flat on the table as if the hidden box had vanished into thin air, only for the box to reappear when the screen returned. Infants fixated significantly longer on the impossible event. Even though the box was completely hidden behind the screen during the rotation, infants knew the box was still there, occupying space, and expected it to physically obstruct the screen’s path. Piaget’s manual search failure was not an ontological void, but a developmental dissociation between early-maturing conceptual representations and late-maturing frontal-motor execution networks.
4.3 The Tunnel Effect, Occlusion, and Discontinuous Trajectories
To further isolate the mechanical precision of the continuity principle, Spelke and her colleagues devised sophisticated multi-occluder paradigms designed to test the limits of what cognitive psychologists refer to as the “tunnel effect.” In a classic design, two separate, wide opaque screens were positioned side by side on a stage, with a clearly visible, empty spatial gap separating them. A small toy or cylinder was set in motion from the extreme left. The object traveled smoothly across the open stage, disappeared behind the first screen, emerged into the open spatial gap between the screens, disappeared behind the second screen, and finally emerged from the extreme right side.
Once infants were familiarized with this continuous, normal trajectory, Spelke introduced an impossible, discontinuous condition. The object was set in motion from the left and disappeared behind the first screen. However, instead of passing through the center gap, the object never appeared in the open space between the two screens; nevertheless, after a short temporal pause, an identical object emerged from behind the second screen and rolled to the right. To an adult observer, this event is mechanically impossible unless two separate objects are present, hidden behind each screen, or unless the object underwent an act of spatial teleportation.
Infants as young as four months of age reacted to this discontinuous display with prolonged looking times, registering the violation immediately. They recognized that a single physical object cannot travel from point A to point B without traversing the intervening visible space. If the intervening space is unobstructed and empty, the object must be visible during its transit. When researchers later lifted the screens to reveal only a single object inside the apparatus, infants registered deep surprise; conversely, if the display had originally featured two objects moving in sequence, no surprise was observed. This proved that infants use the continuity of spatiotemporal paths as a rigorous mathematical tool to individuate objects: one continuous path equals one physical object; a broken, discontinuous path demands the instantiation of two distinct object tokens in working memory.
5. The Principle of Solidity: Mutual Exclusivity in Spatial Occupation
5.1 The Impenetrability Axiom in Core Knowledge
The third core physical constraint identified by Elizabeth Spelke is the Principle of Solidity. Rooted in classical Newtonian mechanics and the fundamental physics of matter, the solidity principle asserts that two solid bodies cannot occupy the same region of space at the same temporal instance. Under this impenetrability axiom, physical objects are bounded by rigid or semi-rigid surfaces that resist spatial intrusion by other solid objects. When two solid objects are placed on intersecting trajectories, they must either undergo an elastic or inelastic mechanical collision, deflect, or come to a complete rest; one solid object cannot pass effortlessly through the interior volume of another.
In the infant’s conceptual architecture, the solidity principle relies upon an innate operational distinction between penetrable media and impenetrable surfaces. Infants readily perceive that their own limbs or moving toys can pass effortlessly through air, through beams of ambient light, or through liquids such as water. In sharp contrast, when encountering solid boundaries—such as floors, wooden shelves, walls, or solid geometric blocks—the core physics engine imposes an absolute barrier to spatial interpenetration. Solidity is not perceived merely as a visual texture or a tactile density; it is an abstract geometric rule regarding the mutual exclusivity of spatial occupancy.
This principle ensures that the infant views the physical environment as a landscape of mechanical boundaries and non-negotiable obstacles. The infant mind assumes boundary rigidity by default. An object is represented as occupying a three-dimensional volume that is fully filled with matter, precluding any other volumetric mass from entering that spatial domain unless the structural boundaries of the host object are catastrophically fractured or deformed. Solidity thus acts as a protective spatial shield in the mind of the infant, orchestrating accurate expectations regarding mechanical interactions, occlusions, and physical blockages.
5.2 Violation Paradigms: Solid Objects Passing Through Barriers
The definitive experimental validation of the solidity principle was achieved in a classic 1992 study conducted by Elizabeth Spelke, Karen Breinlinger, Kathy Macomber, and Kirsten Jacobson, titled “Origins of Knowledge.” This landmark paper deployed a series of meticulously controlled violation-of-expectation experiments to assess whether three- and four-month-old infants understand that falling or rolling objects cannot penetrate solid horizontal or vertical barriers.
In the foundational vertical paradigm, infants watched a small blue ball drop from the top of an apparatus, falling vertically behind an opaque screen until it came to rest on the visible bottom floor of the apparatus. Once habituated to this event, a solid wooden shelf was introduced, mounted horizontally across the middle of the apparatus, clearly above the bottom floor. The opaque screen was lowered to cover the middle portion of the trajectory—occluding the shelf and the space immediately above and below it—while leaving the top drop-point and the bottom floor fully visible. The ball was dropped behind the screen.
Infants were then shown two outcomes when the screen was raised. In the possible (consistent) outcome, the ball was revealed resting peacefully on top of the solid horizontal shelf—its downward trajectory arrested by the impenetrable barrier. In the impossible (inconsistent) outcome, the ball was revealed resting on the bottom floor of the apparatus, directly beneath the shelf. For the ball to have reached this lower position, it would have had to pass directly through the solid, impenetrable wooden shelf. Infants looked significantly longer at the impossible outcome. Despite having seen the ball land on the bottom floor dozens of times during the habituation phase (making that position visually familiar), they rejected that location when a solid barrier blocked the path, proving that their looking behavior was governed by a rich, internal mental representation of physical solidity and mechanical obstruction.
5.3 Solidity Processing Versus Perceptual Feature Analysis
A critical question pursued by developmental cognitive scientists was whether the infant’s response to solidity violations was merely a shallow visual heuristic—such as tracking visual line intersections or expecting an object to disappear when covered—or whether it represented genuine, abstract reasoning about the internal, volumetric properties of matter. Extensive follow-up investigations confirmed that infants reason about internal solid properties that transcend simple optical surface features.
In cross-modal transfer studies combining visual and tactile modalities, researchers demonstrated that the infant’s understanding of solidity is not confined to the visual cortex. For example, when infants are allowed to explore a solid object haptically in the dark or without visual feedback, and are subsequently presented with visual events where that object interacts with other solid bodies, they project identical solidity constraints. The cognitive token instantiated by tactile exploration immediately inherits the mathematical rule of spatial impenetrability. The infant’s mind does not process solidity as an isolated optical effect, but as an amodal property of physical matter that applies universally whether an object is seen, touched, or mentally tracked behind an occluder.
Furthermore, when researchers manipulated the physical state of the barrier, substituting solid boards for penetrable screens, perforated grids, or liquid surfaces, infants systematically altered their expectations. When a ball was dropped above a container filled with water or a frame filled with thin, breakable paper or loose cloth, infants did not show prolonged looking when the falling object appeared at the bottom. Their core physical engine accurately recognized that solidity is a unique mechanical property of solid, bounded masses, rather than a universal property of all spatial surfaces. When interacting with genuine solids, the impenetrability axiom remained absolute, confirming that infant physics is organized around profound mechanical principles rather than superficial perceptual associations.
6. The Principle of Contact: Collision Mechanics and Distant Interactions
6.1 The No-Action-at-a-Distance Constraint
The fourth foundational pillar in Elizabeth Spelke’s taxonomy of core physics is the Principle of Contact. Derived directly from classical contact mechanics, this principle dictates that inanimate objects move exclusively when subjected to direct physical contact; they cannot move themselves, nor can they interact across empty space without a physical intermediary. The contact principle imposes an absolute “no-action-at-a-distance” constraint upon the behavior of inanimate physical bodies, operating as the psychological root of Newtonian force dynamics in human ontogeny.
In the experiential world of the infant, objects such as balls, chairs, blocks, and rocks do not spontaneously accelerate, hover, or launch across rooms of their own volition. Unless acted upon by an applied mechanical force—an external push, a pull, a collision, or a drop—inanimate matter remains static. The infant core knowledge system codifies this mechanical law into a rigorous inferential rule: if an inanimate object transitions from a state of rest to a state of motion, an external contact event must have preceded that transition. Conversely, if an object collides with another, a transmission of kinetic force is expected to occur immediately.
The contact principle serves as a cognitive defense mechanism against magical or teleological thinking in the physical domain. In the absence of this innate constraint, an infant would have no mathematical or logical basis for attributing physical causation, leaving them vulnerable to interpreting every random environmental displacement as an arbitrary, supernatural, or causally disconnected event. By imposing the contact requirement, the human infant instinctively seeks out the mechanical source of any observed kinematic displacement, laying the empirical groundwork for causal inference, tool use, and mechanical problem-solving.
6.2 Michottean Launching and Causal Perception in Infancy
The empirical foundation for the study of causal contact mechanics originated in the classic twentieth-century work of Belgian psychologist Albert Michotte. Michotte developed visual kinetic displays where a colored disc (Object A) moved across a horizontal path toward a stationary colored disc (Object B). When Object A reached Object B, it stopped instantly, and Object B immediately began moving along the exact same trajectory at the same or slightly reduced speed. Adults viewing this display do not report seeing two independent kinematic movements (A moving and stopping; B starting and moving); instead, they experience an irresistible, direct perceptual illusion of causal launching—they see Object A strike Object B, physically pushing it into motion.
Elizabeth Spelke, working in conjunction with researchers such as Alan Leslie, adapted Michottean collision paradigms to preverbal infants using habituation-dishabituation protocols. Leslie and Spelke habituated infants to either a direct causal launching sequence (continuous motion with physical contact) or an impossible, non-contact launching sequence (where Object A stopped short of Object B by several centimeters, yet Object B immediately launched into motion across empty space, exhibiting “action-at-a-distance”).
Following habituation, the researchers reversed the direction of the interaction, or manipulated the temporal and spatial parameters. The results revealed that infants as young as six months of age—and under specific paradigms, as young as four months—are acutely sensitive to the spatial and temporal continuity of collisions. If Object A halts even a fraction of a second before touching Object B, or if Object B delays its launch by a few hundred milliseconds, infants do not perceive a causal collision. They perceive two disconnected, independent events. When infants were habituated to causal launching and subsequently tested with non-causal events, their looking times escalated dramatically. Infants do not learn causal mechanics through years of linguistic training or formal education; they possess an innate, low-level perceptual mechanism that processes direct physical contact as the necessary catalyst for inanimate physical motion.
6.3 The Inanimate-Animate Bifurcation in Contact Mechanics
Crucially, the Principle of Contact does not apply universally across all observed entities in the infant’s ecology; instead, it is restricted to the domain of inanimate physical objects. One of the most brilliant demonstrations of Spelke’s domain-specific cognitive model is the clear empirical bifurcation between core physics and core psychology: infants recognize that human beings, animals, and intentional agents are exempt from the no-action-at-a-distance constraint.
In a series of landmark studies, researchers presented infants with identical kinematic interactions executed by either inanimate objects (e.g., wooden blocks, cardboard cylinders) or animate human agents (e.g., people walking across a room). When a wooden block stopped two feet away from another wooden block, and the second block mysteriously launched into motion across the floor, infants exhibited profound violation-of-expectation looking responses—an action-at-a-distance impossibility had occurred. However, when a human being walked toward another human being, stopped two feet away, and the second human turned and walked away, the infants showed zero surprise.
Infants deploy an early-emerging computational filter that segregates the world into entities governed by mechanical force and entities governed by internal, intentional agency. Animate agents possess internal energy sources; they are self-propelled, capable of spontaneous autonomous motion, and can communicate across empty spatial divides via vocalizations, gestures, or visual gazes. Inanimate objects lack internal energy; they are mechanically inert, strictly bounded by the contact principle. If a cylinder moves on its own without contact, infants actively search the scene for a hidden wire, an incline, or a concealed hand; if a person moves, no such mechanical antecedent is expected. This sharp bifurcation confirms that core physics is not a crude, indiscriminate sensory bias, but an ontologically sophisticated module designed specifically to track inanimate matter.
7. Gravity and Inertia: The Limits and Hierarchies of Early Physical Intuition
7.1 Asymmetry in Core Physics: Early Principles Versus Late-Developing Laws
While the principles of cohesion, continuity, solidity, and contact operate with stunning robustness in the earliest months of human infancy, Elizabeth Spelke’s empirical inquiries revealed a striking, asymmetrical revelation: infants do not possess an innate, fully calibrated understanding of gravity and inertia. This cognitive asymmetry exposed a fundamental architectural hierarchy within human naive physics.
In studies investigating infant reasoning about falling objects, trajectories, and support systems, researchers such as In-Kyeong Kim and Elizabeth Spelke discovered that young infants under five to six months of age often fail to recognize gravitational violations that seem visually glaring to older children and adults. For example, if an unsupported object is released in mid-air and remains magically suspended without physical support, very young infants frequently fail to express surprise, viewing the static hovering object as unproblematic. In contrast, if that same suspended object were to pass through a solid barrier or disappear into thin air, looking-time violations would spike immediately.
From an evolutionary developmental perspective, this hierarchy is profoundly logical. The principles of cohesion, continuity, and solidity are absolute, non-negotiable geometric truths of macroscopic matter across every terrestrial and aquatic ecosystem: a predator or an obstacle remains continuous, solid, and cohesive regardless of whether it is swimming, climbing, resting, or jumping. Gravity and inertia, conversely, represent complex dynamical force interactions whose accurate quantitative calculation depends heavily upon the physical medium, friction, variable surface gradients, mass, and aerodynamic drag. Consequently, natural selection hardwired the absolute topological constraints of objecthood into the initial state, while leaving the fine-grained ballistic mechanics of gravity and inertia to be calibrated through postnatal sensory-motor calibration and systematic visual observation.
7.2 The Development of Support Relations and Gravitational Expectations
The ontogenetic trajectory of how infants come to master gravitational support relations was comprehensively illuminated by Renée Baillargeon, whose findings integrated seamlessly with Spelke’s broader core physics framework. Baillargeon demonstrated that between three and twelve months of age, infants undergo a rigid, systematic sequence of cognitive stages regarding balance, contact, and support.
This developmental progression unfolds across four distinct, rule-governed phases:
- Phase 1 (Initial Contact Rule, ~3 to 4 months): Infants possess a crude, binary heuristic: an object will remain stable if it makes any physical contact with a supporting surface, regardless of where or how that contact occurs. At this stage, an infant is perfectly satisfied if a block is held against the vertical side of a support platform; as long as physical contact exists, gravity is assumed to be neutralized.
- Phase 2 (Type of Contact Rule, ~4.5 to 5.5 months): The infant revises their internal algorithm, realizing that the location of the contact is critical. Infants now recognize that support requires contact on the top or horizontal surface of the support platform. If a block is released against the vertical side of a platform and hovers, they now look significantly longer, registering a violation. However, they still believe that any minimal contact with the top surface is sufficient.
- Phase 3 (Amount of Contact Rule, ~6.5 to 7.5 months): Infants discover that the proportion of the object’s surface area resting on the support is decisive. If a block is placed on a platform with 90% of its base hanging out over empty space (an unstable, off-center position), six-month-olds fail to predict that it will fall; by seven to eight months, infants calculate the center of mass roughly, expecting an object to topple if more than half its base is unsupported.
- Phase 4 (Proportional Mass and Asymmetry Rule, ~11 to 12 months): Infants finally incorporate complex volumetric geometry and asymmetrical mass distribution, recognizing that an L-shaped or weighted object can topple even if more than half its horizontal base is technically supported by the surface.
This systematic developmental ladder demonstrates how the human mind builds upon its innate initial state. The core framework provides the spatial and mechanical scaffolding (the concept of distinct physical objects and the requirement of contact), while iterative observational experience and manual manipulation calibrate the precise mechanical thresholds of gravity and support.
7.3 Inertial Blindness in Infants and Adults
Even more startling than the late emergence of gravity is the pervasive, enduring human vulnerability to inertial blindness. In classical Newtonian physics, Newton’s First Law of Motion dictates that an object in motion will maintain its velocity and direction in a straight line unless acted upon by an external net force. When an object is dropped from a moving carrier (such as an airplane or an infant’s moving hand), the object retains its forward horizontal velocity, tracing a smooth, curved, forward parabolic arc toward the ground.
Studies conducted by Elizabeth Spelke, Michael McCloskey, and their colleagues revealed that neither infants nor adults intuitively calculate inertial parabolic trajectories accurately. When infants and young children are asked to predict where a ball will land when dropped from a horizontally moving train, they exhibit a persistent, systematic error known as the straight-down gravitational bias. They predict that the moment the object is released from the moving vehicle, it will plummet in a perfectly straight vertical line perpendicular to the floor, completely ignoring the forward inertial momentum inherited from the carrier.
Remarkably, this inertial blindness does not dissolve with adulthood or formal primary education. In classic cognitive studies of “folk physics” or “naive mechanics,” Michael McCloskey presented university physics students with a diagram of a curved, spiral tube resting horizontally on a table. When asked to draw the trajectory a marble would take as it emerged from the end of the curved tube, more than a third of the adults drew a curved, spiral trajectory through empty air, operating under the mistaken, medieval Aristotelian notion of “impetus”—the belief that curved motion imparts an internal, circular curving force within the object. Elizabeth Spelke’s work proved that human core physics is decidedly not Newtonian; it is a specialized, terrestrial ecological heuristic optimized for low-velocity, high-friction, immediate environments. Where core physics lacks explicit hardwiring (as in ballistic inertia), human intuition defaults to systematic mechanical illusions that endure throughout life.
8. Methodological Paradigms: The Mechanics and Critiques of Looking-Time Research
8.1 The Violation-of-Expectation (VoE) Technique
The monumental theoretical edifices constructed by Elizabeth Spelke and modern developmental cognitive science rest almost entirely upon the methodological validity of the Violation-of-Expectation (VoE) paradigm. Given the non-verbal, motorically limited nature of the infant, researchers had to design an experimental apparatus that could transform micro-variations in ocular behavior into rigorous, quantifiable indices of underlying conceptual architecture.
The standard VoE protocol typically operates through two primary variants: habituation-dishabituation designs and familiarization-free violation designs. In a habituation paradigm, the infant is seated in a sound-attenuated, dimly lit testing booth facing an illuminated stage. The infant is exposed to repeated presentations of a baseline physical event (e.g., a ball rolling down a ramp behind a screen). An infrared camera or eye-tracker mounted beneath the stage records the infant’s ocular fixations in real time. Observers, strictly blinded to whether the infant is currently viewing a possible or impossible condition, monitor visual fixations. An automated computer algorithm calculates habituation: once the infant’s looking time across three consecutive trials drops below 50% of their initial fixation duration, the habituation criterion is achieved—the infant is bored, having successfully extracted the predictable physical parameters of the display.
At this juncture, the experimental manipulation is introduced: the test phase. The infant is presented with alternating trials of structurally equivalent events that culminate in either a physically consistent (possible) or physically inconsistent (impossible) state. Epistemologically, the entire paradigm relies upon an elegant premise: human infants are cognitive informavores. They allocate their limited attentional resources toward epistemic violations—events where the incoming sensory input conflicts with the predictive simulation generated by the internal cognitive model. In recent decades, researchers have augmented standard looking-time duration with high-precision pupillometry and corneal reflection eye-tracking. Pupillary dilation, driven by locus coeruleus-norepinephrine system activation, serves as an autonomic, involuntary biological index of cognitive surprise and mental effort, corroborating that elevated looking times represent genuine cognitive shock rather than passive, unthinking staring.
8.2 Perceptual Low-Level Critiques: Bogartz, Haith, and Cashon
Despite its profound impact, the Violation-of-Expectation paradigm has faced ferocious methodological and philosophical critiques. Skeptical developmental psychologists, most notably Richard Bogartz, Marshall Haith, and Cara Cashon, published influential critiques arguing that Spelke, Baillargeon, and other nativists were guilty of over-interpreting looking-time data. Haith famously warned against attributing “rich” conceptual, propositional theories of physics to four-month-old minds when “lean,” low-level sensory explanations could fully account for the observed empirical effects.
The core of the low-level critique revolves around perceptual salience, sensory flicker, and motion artifacts. Bogartz and colleagues argued that impossible test events are often visually more dynamic, contain sharper luminance transitions, feature longer total motion paths, or present more novel visual surface areas than their possible counterparts. For example, in the drawbridge experiment, the impossible 180-degree rotation involved 68 degrees of additional physical movement compared to the 112-degree stopped condition. Did the four-month-old infant look longer at the 180-degree event because they possessed an abstract, Kantian understanding of the solidity and permanence of the hidden box, or did their immature visual cortex simply track the extra visual motion and increased optical flow across their retina?
Elizabeth Spelke and her colleagues responded to these critiques with exceptional experimental rigor, establishing methodological counter-controls that systematically dismantled the low-level sensory alternative. Spelke designed experiments where the impossible event was perceptually identical to the habituation event, or where the impossible condition involved less visual motion than the possible condition. In one classic control for solidity, infants were habituated to a ball resting in an open box; during the test phase, an impossible display featured the ball remaining stationary, while a possible display featured the ball rolling across the screen. Infants systematically looked longer at the stationary impossible event despite its visual simplicity. Furthermore, researchers demonstrated that if the exact same physical displays were modified so that the occlusion was caused by a clear, transparent glass pane (where no conceptual violation occurred), infants showed no elevated looking time, despite the identical optical trajectories. These controls definitively established that looking-time elevations are driven by epistemic violations of abstract physical rules, not low-level optical noise.
8.3 Alternative Methodological Approaches and Corroborating Data
To eliminate lingering methodological skepticism regarding passive looking times, contemporary cognitive scientists have mobilized an advanced battery of convergent methodological tools to corroborate Spelke’s core physical framework. Chief among these innovations is the deployment of infant electroencephalography (EEG) and event-related potentials (ERPs).
When infants view physical violations, neuroimaging reveals distinct electrophysiological signatures in real time. Studies measuring infant EEG during occlusion and solidity violations have documented significant bursts of gamma-band oscillatory activity (around 40 Hz) over parietal and frontal scalp electrodes. In adult cognitive neuroscience, gamma oscillations are the definitive neural signature of active object representation and mental maintenance in working memory. When an object passes behind an occluder, the infant’s gamma oscillations remain sustained throughout the occlusion period; if the object magically disappears or violates solidity upon screen removal, a sharp, transient negative ERP component—the Nc component (an index of attentional orienting and unexpected novelty)—is elicited, followed by a sudden collapse of gamma synchrony. This provides irrefutable, millisecond-by-millisecond neurobiological evidence that the infant brain actively maintains a continuous, internal physical representation of the hidden object.
Furthermore, researchers have embraced anticipatory eye-tracking paradigms. Instead of merely measuring how long an infant stares at an event after it has already occurred, modern high-speed eye-trackers record where the infant looks before an object re-emerges from behind a barrier. When an object rolls toward an opaque screen, infants as young as four to five months of age do not track the screen passively; their gaze sweeps instantly across the barrier, fixating proactively upon the opposite exit portal moments before the object emerges. If a solid barrier is visibly lowered behind the screen, the infant’s anticipatory saccade adjusts, landing precisely at the point of expected collision. Finally, ingenious manual reaching assays—such as assessing infant reaching in complete darkness toward the remembered spatial coordinates of an occluded, glowing, or sound-emitting object—demonstrate that when motor planning demands are radically simplified, the infant’s physical motor output aligns perfectly with their internal Spelkean core knowledge.
9. Core Knowledge Versus Piagetian Constructivism: A Paradigm Shift
9.1 The Deconstruction of the Sensorimotor Stage
The emergence of Elizabeth Spelke’s Core Knowledge Theory precipitated a tectonic paradigm shift in developmental psychology, fundamentally deconstructing Jean Piaget’s venerable Sensorimotor Stage. For decades, the dominant Piagetian architecture maintained that conceptual cognition is the late-emerging byproduct of internalized physical actions. In Piaget’s classic formulation, there is no conceptual thought without prior sensory-motor coordination; abstract understanding of space, time, matter, and causality must be synthesized from scratch through manual grasping, oral exploration, dropping objects, and physical locomotion during the first eighteen to twenty-four months of life.
Spelke’s empirical discoveries completely inverted this classical causality. Rather than conceptual representation being the hard-won product of motor exploration, Spelke demonstrated that conceptual knowledge precedes, guides, and structures motor mastery. Preverbal infants who cannot yet reach, grasp, roll over, or crawl already possess sophisticated, abstract, non-perceptual representations of physical bodies, their boundaries, their permanence, and their mechanical interactions. The infant is not an unthinking sensorimotor machine that slowly manufactures thought through reflex manipulation; the infant is an innate conceptual thinker whose physical actions are guided from the outset by core inductive engines.
This inversion resolved a deep theoretical paradox at the heart of constructivism. If an infant had no antecedent concepts of object boundaries, cohesion, or permanence, manual exploration would be cognitively paralyzing. When an infant grasps a cup, why do they not expect the handle to rip away from the vessel like liquid? When they drop a block onto a table, why are they not shocked that the table does not swallow the block like a shadow? Without innate core physical constraints, an infant’s sensorimotor experiences would be an uninterpretable torrent of unstructured sensory noise. Core knowledge provides the foundational ontological categories that make sensorimotor learning possible in the first place.
9.2 The Competence-Performance Distinction in Cognitive Development
The clash between Spelke’s nativism and Piaget’s constructivism forced developmental psychology to embrace a rigorous epistemological distinction: the Competence-Performance Distinction. Originally articulated by Noam Chomsky in the analysis of linguistic grammar, this distinction separates an organism’s underlying cognitive knowledge (competence) from its ability to deploy that knowledge within real-time, real-world behavioral tasks that impose heavy non-conceptual demands (performance).
The notorious “A-not-B error”—where an infant watches an experimenter hide a toy at Location A several times and retrieves it successfully, but continues to reach toward Location A even after visibly watching the toy hidden at Location B—was historically interpreted by Piaget as proof that the infant believed the object’s physical existence was contingent upon their own manual action (egocentric causality). However, when researchers such as Adele Diamond and Elizabeth Spelke investigated the A-not-B error using non-manual, eye-tracking paradigms, they uncovered a stunning divergence: during the search phase, the infant’s eyes and anticipatory gaze fixate reliably upon Location B (the true location of the object), while their hand automatically reaches back toward Location A.
This profound dissociation proved that the A-not-B error is an executive motor-performance failure, not a conceptual deficit. To manually retrieve the object from Location B, the infant must integrate several immature neuro-developmental systems: they must maintain the spatial coordinates in working memory, suppress a previously rewarded motor habit (prepotent motor inhibition), and coordinate a complex, two-step means-end motor plan (removing the barrier and grasping the toy). The dorsolateral prefrontal cortex, which governs executive working memory and motor inhibition, matures exceptionally slowly throughout the first year of life. When infants fail manual search tasks, they fail because the frontal-striatal motor execution pipeline is overwhelmed, not because the core knowledge of object continuity has evaporated. In looking-time tasks, where frontal-motor demands are eliminated, the infant’s pure conceptual competence shines through with crystalline clarity.
9.3 Representational Formats: Perceptual Schema Versus Propositional Knowledge
The philosophical and cognitive debate over core physics naturally extended into the realm of representational formats: in what cognitive language is core knowledge written? Scholars such as Jean Mandler, Lawrence Barsalou, and Susan Carey engaged deeply with Spelke’s claims, questioning whether the infant’s understanding consists of low-level perceptual schemas, rich propositional knowledge, or something entirely unique.
One prominent alternative to Spelke’s nativist rationalism was Jean Mandler’s Image Schema Theory. Mandler proposed that infants do not possess innate, language-like, propositional axioms about physics; instead, they deploy an early perceptual mechanism called “perceptual analysis” that recodes complex spatial-temporal visual impressions into schematic, abstract spatial primitives or “image schemas” (such as PATH, CONTAINER, LINK, and BLOCKAGE). These image schemas function as an intermediate bridge between raw sensory perception and abstract conceptual thought, providing an analog, topological format for spatial reasoning without requiring an innate language of thought.
Elizabeth Spelke maintained a distinctly conceptual, inferential stance. While acknowledging that core knowledge operates without conscious, verbal awareness, Spelke argued that core physical principles function as true inferential computational systems. They are not merely passive pictorial summaries or visual templates of past experiences; they are generative, deductive engines. Core representations exhibit compositionality and systematicity—the hallmark characteristics of conceptual thought. An infant can apply the principles of cohesion and solidity to entirely novel, bizarre shapes that they have never seen before in their evolutionary or postnatal history. The representational format of core physics is an abstract, domain-specific, symbolic-geometric code that operates upstream of motor control and downstream of raw perceptual transduction.
10. Evolutionary and Comparative Foundations of Physical Cognition
10.1 Phylogenetic Heritage: Physical Reasoning in Non-Human Primates
If the core knowledge of physics represents an evolutionary adaptation hardwired into the human neonate, it must not be a biological anomaly restricted to Homo sapiens. Evolutionary developmental biology demands that such foundational survival mechanisms exhibit profound phylogenetic continuity across our closest living evolutionary relatives. Over the past three decades, extensive comparative cognitive research has established that non-human primates share the identical core physical architecture documented by Elizabeth Spelke in human infants.
Comparative psychologists, including Marc Hauser, Laurie Santos, and Josep Call, deployed Spelkean violation-of-expectation and looking-time apparatuses to test rhesus macaques (Macaca mulatta), cotton-top tamarins (Saguinus oedipus), and chimpanzees (Pan troglodytes). When tested with identical drawbridge apparatuses, multi-occluder continuity tunnels, and horizontal solidity shelves, non-human primates exhibited behavioral responses indistinguishable from human infants. Rhesus macaques look significantly longer when an apple passes through a solid barrier, when an object teleports across an unoccluded spatial gap, or when a cohesive entity spontaneously fractures into fragments without external force.
Furthermore, primates demonstrate the exact same hierarchical cognitive boundaries observed in human ontogeny: while they master continuity, cohesion, and solidity effortlessly, they struggle with the precise mechanics of balance, center of mass, and ballistic inertia. This deep evolutionary conservation demonstrates that the core knowledge of physics emerged hundreds of millions of years before the advent of human language, symbolic culture, or tool manufacture. It represents a primitive, highly conserved vertebrate neural architecture designed to navigate the macroscopic physical ecology of planet Earth.
10.2 Physical Knowledge in Avian and Other Non-Primate Models
The evolutionary reach of core physical cognition extends far beyond our primate cousins. Astonishing empirical discoveries within avian ethology and comparative cognition have revealed that birds—specifically corvids (crows, ravens, Eurasian jays) and gallinaceous species (domestic chicks, Gallus gallus domesticus)—possess fully operational core physical reasoning systems that exhibit stunning convergent evolution with mammalian architecture.
In groundbreaking research led by Giorgio Vallortigara and his colleagues, newly hatched domestic chicks were tested within hours of hatching, completely ruling out prior visual experience, trial-and-error learning, or motor conditioning. Using filial imprinting protocols—where newborn chicks naturally imprint upon a moving artificial geometric object—researchers placed the imprinted object behind an occluder, manipulated its trajectory, or simulated solidity and continuity violations across complex multi-barrier mazes. Without a single day of prior physical practice, newborn chicks spontaneously tracked the unperceived imprinted object behind opaque barriers, calculated its continuous spatiotemporal trajectory, and anticipated its emergence with millisecond precision.
Similarly, research on New Caledonian crows by Nicola Clayton and Alex Taylor demonstrated that corvids possess sophisticated, causal understanding of water displacement (the classic Aesop’s fable paradigm), solidity, and contact mechanics during spontaneous tool modification. Crows intentionally select heavy, solid objects over hollow, buoyant, or liquid objects to drop into water tubes to raise the water level and retrieve floating food. Because the avian brain lacks a laminated mammalian neocortex—operating instead through a densely packed, non-layered pallial structure known as the dorsal ventricular ridge (DVR)—these findings prove that core physical knowledge does not depend on a uniquely human or primate cortical architecture. Rather, core physics is a universal computational solution that natural selection has independently engineered across divergent vertebrate neuroanatomical platforms.
10.3 Evolutionary Adaptiveness and Ecological Validity
Why did natural selection hardwire mechanical constraints so deeply into the biological architecture of diverse animal species? The answer lies in the harsh mathematical logic of evolutionary survival and ecological validities. In the wild, trial-and-error associative learning is an evolutionary luxury that fragile, vulnerable offspring simply cannot afford.
Consider the ecological stakes: if a newborn primate or avian chick were required to learn through unguided, empirical reinforcement that solid objects cannot be walked through, that physical support is necessary to avoid catastrophic falls down ravines, or that an approaching predator does not cease to exist when it momentarily slips behind a boulder, that organism would not survive its first week of life. The biological cost of an erroneous physical inference is catastrophic, immediate mortality. Conversely, the metabolic cost of hardwiring a small set of foundational domain-specific inductive priors into embryonic neurodevelopment is exceptionally modest.
This ecological imperative explains the biological trade-off between rigid core modularity and open-ended plastic intelligence. Core physics provides an unbreakable, reliable bedrock of spatial and mechanical survival rules that function flawlessly from day one, in zero-shot environments. Upon this immutable modular foundation, evolution subsequently layered more plastic, late-maturing cortical associative systems capable of fine-tuning support calculations, mastering local frictional environments, and ultimately inventing formal, symbolic scientific tools. Core knowledge represents nature’s ultimate risk-mitigation strategy: an unyielding cognitive scaffold guaranteeing that whatever else the mind learns, it never loses its grip on physical reality.
11. Ontogenetic Development: Extending and Transcending Core Physics
11.1 The Conceptual Change Framework: Enrichment Versus Restructuring
If human infants begin life equipped with an innate, modular system of core physics, how do they eventually transform into adults capable of understanding Newtonian mechanics, orbital astrophysics, or quantum theory? This question ignited one of the most celebrated intellectual debates in modern cognitive science, pitting Elizabeth Spelke against her long-time collaborator and friend, Susan Carey, over the nature of Conceptual Change.
Elizabeth Spelke championed an enrichment perspective. Spelke argued that core knowledge systems are never rewritten, abandoned, or fundamentally dismantled over ontogenetic development. Instead, they remain permanently intact within the adult cognitive architecture, operating as continuous, domain-specific modules. Development, under Spelke’s view, consists of conceptual enrichment: novel concepts are added around the core modules, and new combinatorial systems (principally natural language) construct cognitive bridges connecting previously isolated core domains. When a child learns formal scientific physics, they do not erase their core knowledge of solidity or contact; rather, they learn an alternative, culturally invented explicit theoretical framework that operates alongside their innate intuitions.
Susan Carey countered with a radical restructuring perspective, heavily influenced by Thomas Kuhn’s philosophy of scientific revolutions. Carey argued that cognitive development involves profound “Kuhnian” paradigm shifts in which early cognitive frameworks are structurally transformed and reorganized. In Carey’s model of “bootstrapping,” the child encounters physical phenomena that cannot be adequately explained by their initial core theories (such as the behavior of gases, biological growth, thermal expansion, or non-cohesive matter). Through the use of external cultural symbols, mental modeling, and metaphorical analogies, the child undergoes a qualitative conceptual revolution, fundamentally redefining foundational ontological primitives (e.g., differentiating between “weight” and “density,” or between “matter” and “occupying space”). While Carey acknowledged the reality of Spelkean initial states, she insisted that human development possess the power to genuinely transcend and restructure the initial state.
11.2 The Role of Natural Language as a Cognitive Combinatorial Engine
How does the human mind overcome the strict modular encapsulation of core knowledge systems to produce flexible, creative, unified adult thought? Elizabeth Spelke formulated a revolutionary hypothesis: Natural Language serves as the universal cognitive combinatorial engine that breaks down the walls separating distinct core modules.
In non-human animals and preverbal infants, core modules operate in splendid isolation:
- Core Physics tracks individual bounded objects, collisions, and spatial occlusions.
- Core Geometry computes spatial reorientation based on the large-scale metric shape of the surrounding environment.
- Core Number tracks small, exact sets of items (1, 2, 3) through parallel indexing, or calculates large approximate quantities via the Approximate Number System (ANS).
- Core Psychology attributes goals, perceptions, and desires to animate agents.
In a preverbal infant or a primate, these systems cannot cross-communicate. A classic example emerges from Spelke’s spatial reorientation experiments. When an infant is disoriented in a rectangular room with one blue wall, they can easily use the core geometric module to locate a hidden object in diagonally opposite corners (utilizing room geometry), and they can use core physics to track an object; however, they cannot combine these domains to form the conjunctive thought: “The toy is at the corner to the left of the short blue wall.” They possess the geometric representation (short wall) and the featural/physical representation (blue surface), but lack the cognitive glue to combine them.
Natural language changes everything. Around the age of four to six years, as children acquire spatial prepositions (“to the left of”), lexical color terms (“blue”), and syntactically governed noun-modifier combinations, they suddenly master these cross-modular integration tasks. Spelke posited that the compositional syntax of natural language provides a domain-general combinatorial medium. Language allows predicates from core physics, geometric mapping, and numerical tracking to be merged into novel, composite conceptual propositions. Natural language is the evolutionary catalyst that allows humanity to transcend core modularity and construct science, art, and philosophy.
11.3 The Survival of Core Physics in Adult Intuitions and Folk Mechanics
A central tenet of Spelke’s framework is that core knowledge is never erased; it persists indefinitely beneath our civilized, educated cognitive surfaces. Despite years of formal secondary and university instruction in Newtonian mechanics, thermodynamics, and Einsteinian relativity, adult human beings consistently default to ancient Spelkean intuitions whenever they are required to make rapid, intuitive physical judgments under time pressure.
This reality is vividly illustrated in cognitive psychology experiments on adult folk mechanics. When adult university graduates—many of whom have completed college-level physics courses—are placed in computer simulation environments and tasked with predicting the real-time flight of dropped objects, the trajectories of cut swinging pendulums, or the path of water spraying from a curved garden hose, their explicit Newtonian training frequently evaporates. They predict that a swinging pendulum, when cut at the apex of its swing, will fly straight outward as if imbued with centrifugal momentum; they predict that an object dropped by a running person falls backward or straight down; they predict that heavy objects fall inherently faster than light objects in a vacuum.
Cognitive science reconciles this persistent divergence through dual-processing models of cognition (System 1 versus System 2). System 1 consists of the fast, automatic, unconscious operations of our evolutionary core knowledge modules. When you throw a ball, catch a falling glass, or duck beneath an incoming projectile, System 1 core physics executes lightning-fast kinematic simulations to guide motor survival. System 2, conversely, represents the slow, deliberate, culturally acquired, symbolic machinery of formal scientific mathematics. While System 2 can calculate Newton’s equations on a chalkboard, System 1 core physics remains the default operating system running in the background of every human mind from infancy to death.
12. Contemporary Implications: Developmental Neuroscience and Artificial Intelligence
12.1 Neural Mechanisms of Core Object Knowledge
In the twenty-first century, the insights of Elizabeth Spelke have converged powerfully with the frontiers of developmental cognitive neuroscience. Using advanced, non-invasive neuroimaging modalities—such as high-density functional Near-Infrared Spectroscopy (fNIRS) and infant functional Magnetic Resonance Imaging (fMRI)—neuroscientists have begun mapping the explicit anatomical substrates of the core intuitive physics engine in the human brain.
Studies conducted on infants ranging from three to twelve months of age demonstrate that core physical reasoning is not distributed homogenously throughout the cortex; rather, it activates a specialized, highly integrated frontoparietal-temporal network. When infants witness violations of continuity, cohesion, or solidity, significant focal blood-oxygenation changes are observed in the inferior parietal lobule (IPL), the intraparietal sulcus (IPS), and regions of the dorsal premotor cortex. These parietal areas are precisely homologous to the brain networks identified by Nancy Kanwisher and Jason Fischer in adults as hosting the human “intuitive physics engine.”
This dorsal stream-dominant network functions as a dedicated real-time physical simulator. It interfaces closely with the temporal visual motion area (MT+/V5) to extract kinematic vectors, while querying frontal executive zones to generate forward predictive simulations. Long before an infant develops formal motor control, these parietal circuits are already hardwired with specific computational priors regarding spatial occupancy, volumetric boundary constraints, and mechanical momentum. The Spelkean infant mind is mirrored directly in the biological hardware of the human cerebral architecture.
12.2 Computational Models of Intuitive Physics Engines
Within computational cognitive science, Elizabeth Spelke’s empirical principles have inspired a profound theoretical renaissance led by researchers such as Joshua Tenenbaum, Peter Battaglia, and Tomer Ullman at MIT and Harvard. These computational theorists have formalized Spelke’s core physics into what is now widely known as the Intuitive Physics Engine (IPE) model.
The IPE framework posits that the human mind reasons about the physical world using computational mechanisms analogous to modern three-dimensional computer graphics and video game physics engines (such as Unity or Unreal Engine). Rather than relying upon brittle, lookup tables of past visual memories, or running complex differential calculus equations, the human mind constructs a probabilistic, generative physical simulation. When an infant views a tower of unstable wooden blocks, their core physics engine instantiates a coarse, low-resolution 3D volumetric model of the blocks, attributes approximate mass and friction parameters to the items, and runs forward Monte Carlo simulations under the influence of gravity to evaluate whether the structure will topple.
Crucially, Tenenbaum and colleagues emphasize that human intuitive physics is probabilistically approximate, not mathematically exact. The core physics engine does not solve Newtonian mechanics analytically; it runs fast, heuristic mental simulations that incorporate sensory uncertainty and noise. This computational formalization mathematically explains both the breathtaking triumphs and the systematic blind spots of infant and adult physical cognition: the infant can predict with exquisite speed whether a falling object will be stopped by a shelf, yet struggles to trace the precise inertial trajectory of an accelerated parabolic arc.
12.3 Artificial Intelligence and the Quest for Machine Common Sense
The field of modern Artificial Intelligence (AI) has encountered a monumental theoretical barrier directly connected to the work of Elizabeth Spelke. Despite the breathtaking triumphs of contemporary deep learning—from multi-billion-parameter Large Language Models (LLMs) to massive generative computer vision systems—artificial intelligence remains notoriously brittle, fragile, and profoundly lacking in basic physical common sense.
State-of-the-art deep neural networks trained on petabytes of internet video can generate photorealistic imagery, yet they routinely produce video sequences where solid objects phase magically through one another, where limbs spontaneously disintegrate (violating cohesion), where entities teleport across backgrounds (violating continuity), or where unsupported objects hover arbitrarily in mid-air. Deep learning systems operate primarily through brute-force pattern recognition, surface-level statistical correlations, and pixel-level predictive interpolations. Because they lack innate architectural constraints, they possess no internal ontological model of an “object” as a bounded, enduring, solid entity that persists across time and space.
To overcome this catastrophic bottleneck, AI pioneers such as Yann LeCun, Demis Hassabis, and Josh Tenenbaum have explicitly argued that artificial general intelligence (AGI) cannot be achieved merely by scaling up domain-general connectionist models on more data. Instead, AI must return to the foundational lessons of Elizabeth Spelke: systems require innate inductive biases. Computer scientists are now actively designing hybrid neural architectures that bake Spelkean principles—cohesion, continuity, solidity, and contact—directly into the neural network’s foundational code. By equipping autonomous robots and artificial visual systems with an explicit “Spelkean prior,” machines are finally beginning to acquire the robust, zero-shot physical common sense that a four-month-old human infant deploys effortlessly in the cradle.
Conclusion: The Architecture of the Infant Mind and its Lasting Legacy
Elizabeth Spelke’s monumental lifelong inquiry into the infant mind has fundamentally redefined our understanding of human nature, cognition, and epistemology. By penetrating the historical barrier of the infant motor deficit, Spelke overturned centuries of radical empiricism and behaviorist dogma, revealing that the human neonate does not enter existence in a chaotic sensory fog, but arrives armed with a sophisticated, evolutionary inheritance of core physical knowledge.
Through the foundational principles of cohesion, continuity, solidity, and contact, the infant mind carves dynamic reality into an ordered, predictable universe populated by enduring, bounded three-dimensional objects. These core mechanisms operate as autonomous, domain-specific inference engines, providing the indispensable cognitive scaffolding upon which all later learning, sensorimotor mastery, tool invention, natural language acquisition, and formal scientific theorizing are erected. Core knowledge is neither a transient developmental phase nor an imperfect illusion; it is the permanent, biologically conserved computational foundation of human intelligence.
As cognitive science accelerates into the twenty-first century, spanning the frontiers of developmental neuroimaging, evolutionary comparative biology, and artificial intelligence, the profound vision of Elizabeth Spelke continues to guide the quest to decipher the mind. In demonstrating that our deepest concepts of space, time, matter, and cause are written into the very initial state of human biology, Spelke achieved one of the most enduring intellectual syntheses in the history of cognitive science—proving that long before we can speak, reach, or consciously reflect, we are already, at our core, intuitive physicists navigating the architecture of the cosmos.
References
- Baillargeon, R. (1987). Object permanence in 3½- and 4½-month-old infants. Developmental Psychology, 23(5), 655–664. https://doi.org/10.1037/0012-1649.23.5.655
- Battaglia, P. W., Hamrick, J. B., & Tenenbaum, J. B. (2013). Simulation as an engine of physical scene understanding. Proceedings of the National Academy of Sciences, 110(45), 18327–18332. https://doi.org/10.1073/pnas.1306572110
- Bogartz, R. S., Shinskey, J. L., & Schilling, T. H. (2000). Object permanence in five-and-a-half-month-old infants? Infancy, 1(4), 403–428. https://doi.org/10.1207/S15327078IN0104_3
- Carey, S. (2009). The Origin of Concepts. Oxford University Press. https://doi.org/10.1093/acprof:oso/9780195367638.001.0001
- Carey, S., & Spelke, E. (1994). Domain-specific knowledge and conceptual change. In L. A. Hirschfeld & S. A. Gelman (Eds.), Mapping the Mind: Domain Specificity in Cognition and Culture (pp. 169–200). Cambridge University Press. https://doi.org/10.1017/CBO9780511624513.008
- Diamond, A. (1985). Development of the ability to use recall to guide action, as indicated by infants’ performance on AB. Child Development, 56(4), 868–883. https://doi.org/10.2307/1130100
- Fischer, J., Mikhael, J. G., Tenenbaum, J. B., & Kanwisher, N. (2016). Functional neuroanatomy of intuitive physical inference. Proceedings of the National Academy of Sciences, 113(34), E5072–E5081. https://doi.org/10.1073/pnas.1610344113
- Haith, M. M. (1998). Who put the cog in infant cognition? Is rich interpretation too costly? Infant Behavior and Development, 21(2), 167–179. https://doi.org/10.1016/S0163-6383(98)90001-7
- Kellman, P. J., & Spelke, E. S. (1983). Perception of partly occluded objects in infancy. Cognitive Psychology, 15(4), 483–524. https://doi.org/10.1016/0010-0285(83)90017-8
- Kim, I. K., & Spelke, E. S. (1992). Infants’ sensitivity to gravity and inertia in physical events. Cognition, 44(3), 315–330. https://doi.org/10.1016/0010-0277(92)90002-L
- Leslie, A. M. (1984). Spatiotemporal continuity and the perception of causality in infants. Perception, 13(3), 287–305. https://doi.org/10.1068/p130287
- Mandler, J. M. (1992). How to build a baby: II. Conceptual primitives. Psychological Review, 99(4), 587–604. https://doi.org/10.1037/0033-295X.99.4.587
- McCloskey, M. (1983). Intuitive physics. Scientific American, 248(4), 122–130. https://doi.org/10.1038/scientificamerican0483-122
- Michotte, A. (1963). The Perception of Causality. Basic Books. https://psycnet.apa.org/record/1963-08573-000
- Piaget, J. (1954). The Construction of Reality in the Child. Basic Books. https://doi.org/10.1037/11168-000
- Pylyshyn, Z. W. (1989). The role of location indexes in spatial perception: A sketch of the FINST spatial-index model. Cognition, 32(1), 65–97. https://doi.org/10.1016/0010-0277(89)90014-0
- Santos, L. R., & Hauser, M. D. (2002). A non-human primate’s understanding of solidity: Dissociations between visual and manual search tasks. Cognitive Psychology, 44(4), 400–435. https://doi.org/10.1006/cogp.2001.0776
- Spelke, E. S. (1990). Principles of object perception. Cognitive Science, 14(1), 29–56. https://doi.org/10.1207/s15516709cog1401_3
- Spelke, E. S. (1994). Initial knowledge: Six suggestions. Cognition, 50(1–3), 431–445. https://doi.org/10.1016/0010-0277(94)90039-D
- Spelke, E. S., Breinlinger, K., Macomber, J., & Jacobson, K. (1992). Origins of knowledge. Psychological Review, 99(4), 605–632. https://doi.org/10.1037/0033-295X.99.4.605
- Spelke, E. S., & Kinzler, K. D. (2007). Core knowledge. Developmental Science, 10(1), 89–96. https://doi.org/10.1111/j.1467-7687.2007.00569.x
- Vallortigara, G. (2012). Core knowledge of physical objects in the young chick. In The Oxford Handbook of Comparative Evolutionary Psychology (pp. 87–104). Oxford University Press. https://doi.org/10.1093/oxfordhb/9780195394153.013.0006
- Xu, F., & Carey, S. (1996). Infants’ metaphysics: The case of numerical identity. Cognitive Psychology, 30(2), 111–153. https://doi.org/10.1006/cogp.1996.0005