The human capacity to navigate complex physical landscapes, comprehend abstract diagrams, and conceptualize relational structures hinges upon internal cognitive representations that mediate between sensory input and purposeful action. For decades, cognitive science operated under the assumption that internal spatial representations mirror Euclidean cartography—an internal, continuous, and metrically precise coordinate grid often characterized as a “cognitive map.” However, an extensive body of empirical research led by cognitive psychologist Barbara Tversky has fundamentally dismantled this metric view. Tversky demonstrated that internal representations of space are not rigid, photorealistic, or uniformly calibrated cartographic artifacts. Instead, they operate as cognitive collages and structured spatial mental models: coarse, qualitative, fragmented, and hierarchically organized schemas deeply shaped by functional human goals, bodily asymmetries, and perceptual heuristics.
Tversky’s Spatial Mental Models Framework posits that space serves as an overarching, primary scaffolding not only for physical locomotion and perceptual parsing, but for human conceptualization as a whole. Rather than computing continuous geometric coordinates, the human mind segments environments into meaningful figures, anchors them to salient reference objects, and encodes relationships using categorical spatial primitives. These mental configurations are multi-tiered, dynamic, and constructed on the fly to satisfy communicative, navigational, and problem-solving demands. By investigating how humans read maps, recount narratives, interpret diagrams, and coordinate bodily actions, Tversky unveiled a systematic architecture of mind where spatial perception, embodied physical interaction, and abstract thought converge into an integrated semiotic system.
This comprehensive treatise examines the theoretical foundations, structural mechanics, empirical paradigms, and practical applications of Barbara Tversky’s spatial mental models framework. Beginning with its historical departure from early metric map theories and classical mental imagery, the inquiry proceeds through the anatomical and environmental asymmetries that govern human spatial frameworks, the systematic distortions and cognitive heuristics that govern geographic reasoning, and the profound role of embodied action and external visuospatial inscriptions. Finally, it analyzes the contemporary neurocognitive validation of Tversky’s principles, illustrating how her functionalist, qualitative paradigm continues to reshape cognitive psychology, artificial intelligence, urban design, and human-computer interaction.
1. Introduction to Barbara Tversky’s Spatial Mental Models Framework
1.1 Historical Context within Cognitive Science and Spatial Cognition
The scientific conceptualization of internal spatial representation underwent a radical paradigm shift over the course of the twentieth century. Early twentieth-century behaviorism systematically eschewed internal mental states, conceptualizing navigation as a chain of stimulus-response pairings where motor habits were reinforced through environmental rewards. This reductionist view was famously challenged by Edward C. Tolman (1948), whose seminal experiments with rodents in complex mazes revealed the existence of latent learning and unreinforced route optimization. Tolman introduced the foundational concept of the “cognitive map,” positing that organisms construct a comprehensive, field-like internal representation of environmental layouts that enables flexible shortcutting, detour behaviors, and navigational problem-solving. Tolman’s work initiated a cognitive revolution in spatial psychology, asserting that internal representations play a central causal role in environmental interaction.
Despite Tolman’s breakthrough, early spatial cognition theorists frequently over-extended the cartographic metaphor, conceptualizing the cognitive map as a literal, metric internal mirror of Euclidean space. Under this classical assumption, internal spatial representations were presumed to preserve geometric fidelity, linear distances, invariant angular relationships, and continuous Cartesian coordinates. However, emerging research throughout the 1960s and 1970s in schema theory—championed by thinkers such as Frederic Bartlett and subsequently integrated into modern cognitive science by Ulric Neisser—began to expose the flaws of treating mental representations as static, photographic replicas of the physical world. Schemas were understood to be active, reconstructive cognitive organizations that interpret, abstract, and regularize sensory input according to semantic expectations and prior structural knowledge.
The transition from rigid metric cartography to flexible mental models was further catalyzed by Philip Johnson-Laird’s mental models framework (1983). Johnson-Laird argued that human reasoning does not rely on formal syntactic logic or complete propositional inventories, but rather on structural analogues of situations constructed dynamically within working memory. Influenced by these developments, Barbara Tversky recognized that spatial representations are neither metric blueprints nor simple propositional strings. She synthesized the reconstructive nature of cognitive schemas with the relational flexibility of mental models. In doing so, Tversky demonstrated that spatial cognition prioritizes qualitative relationships, behavioral utility, and communicative function over rigid Euclidean geometry, providing a theoretical foundation that resolved decades of paradoxes surrounding human spatial errors and navigational biases.
1.2 The Foundational Premises of Tversky’s Spatial Paradigm
Tversky’s spatial mental models paradigm is grounded in the functionalist premise that human cognition evolved to serve embodied action within physical environments. Physical space is not an inert backdrop across which sensory data is passively registered; rather, it is an overarching organizational structure for thought, perception, and long-term memory. According to Tversky, the mind does not passively reflect objective physical geography. Instead, human conceptualization of space is intrinsically functional: it encodes what is ecologically, behaviorally, and socially salient to the organism. Objective physical reality operates according to continuous metric laws, but subjective psychological space operates according to principles of perceptual grouping, categorical segmentation, and actionable relevance.
This functional distinction leads to a fundamental divergence between absolute objective geometry and perceived subjective space. In objective space, distances are strictly symmetric, triangles abide unconditionally by the triangle inequality, and spatial coordinates exist independently of any particular observer. Subjective cognitive space, conversely, is distorted by attentional focus, environmental salience, and behavioral intention. Distances are systematically warped around prominent landmarks, boundaries act as cognitive barriers that amplify subjective separation, and spatial axes are fundamentally asymmetric based on physical embodiment. Evolutionary pressures did not select for an organism capable of running computationally prohibitive matrix transformations to calculate millimeter-accurate coordinates. Rather, natural selection favored an economical, fast, and computationally frugal cognitive architecture capable of identifying navigational goals, avoiding imminent hazards, and tracking interpersonal interactions.
Consequently, Tversky conceptualizes the mind as prioritizing action-oriented spatial representations. Spatial models preserve the topological and categorical essentials necessary to execute movement, communicate directions, and manipulate tools, while discarding superfluous metric noise. The spatial framework is fundamentally parsimonious: it abstracts continuous spatial flow into discrete, manageable cognitive tokens. By treating space as a dynamic, goal-directed mental construct, Tversky’s paradigm bridged the historical divide between ecological psychology and internal computational modeling, establishing space as the foundational substrate upon which both concrete actions and abstract metaphors are built.
1.3 Differentiating Spatial Mental Models from Mental Imagery
A critical contribution of Tversky’s scholarship is the rigorous theoretical and empirical differentiation between spatial mental models and classical mental imagery. Following the intense “imagery debates” of the 1970s and 1980s—largely contested between Stephen Kosslyn’s pictorial analog view and Zenon Pylyshyn’s propositional critique—spatial representations were frequently conflated with visual images. Pictorial views asserted that mental imagery possesses quasi-visual, depictive properties: an internal mental “screen” where continuous, metric, viewer-dependent arrays of visual information are scanned, inspected, and rotated. Tversky challenged this visual reductionism, proving that spatial mental models represent structural, spatial-relational configurations that are distinct from, and cognitively superordinate to, sensory-specific mental images.
The structural contrast between pictorial analog images and spatial mental models resides primarily in modality-independence and abstractness. While mental images are inherently sensory-bound—typically tied to a particular visual point of view, complete with sensory parameters such as color, texture, illumination, and retinal perspective—spatial mental models are abstract, structural networks of spatial relations. Mental models encode entities and the relational vectors connecting them without necessitating pictorial rendering. For instance, an individual can possess a fully operational spatial mental model of a room’s contents, understanding the qualitative proximities and directional orientations of furniture, without visualizing the specific visual appearance, color palette, or photorealistic lighting of the room. Spatial mental models can be constructed with equal facility through direct visual perception, tactile exploration by visually impaired individuals, or linguistic text descriptions.
This modality-independent nature allows spatial mental models to integrate multimodal sensory information into an organized whole. While an analog image struggles to simultaneously maintain visual, auditory, and haptic vectors without sensory interference, a spatial model unifies kinesthetic cues, cartographic exposure, verbal navigation commands, and optical sights into a coherent relational representation. Furthermore, spatial mental models possess superior computational parsimony and efficiency compared to pixelated mental depictions. Generating and maintaining a fine-grained, continuous visual image imposes an enormous cognitive load on working memory resources. Spatial mental models bypass this bottleneck by relying on qualitative spatial relations—such as containment, connectivity, adjacency, and relative order—enabling rapid spatial reasoning and flexible perspective-taking without the computational overhead of rendering continuous visual scenes.
2. Theoretical Foundations: Cognitive Maps versus Cognitive Collages
2.1 The Fallacy of the Unified Metric Cognitive Map
The theoretical premise of the cognitive map long rested upon the assumption of a globally coherent, Euclidean coordinate system residing within human long-term memory. Under this model, every represented entity in an environment is assigned a localized vector within a continuous metric space, allowing an individual to compute straight-line distances and angular bearings between arbitrary points via geometric formulas. Barbara Tversky systematically dismantled this premise by providing extensive empirical evidence of pervasive coordinate discrepancies, metric axiom violations, and intractable cognitive overhead in human spatial judgments across both geographic and micro-environmental scales.
In a series of landmark studies, Tversky demonstrated that human spatial memory systematically violates the fundamental mathematical axioms of metric space: the identity axiom, the symmetry axiom, and the triangle inequality. According to the metric symmetry axiom, the distance from point A to point B must equal the distance from point B to point A ($d(A, B) = d(B, A)$). However, human directional and distance estimations routinely violate this principle. Prominent, salient reference points—such as major landmarks or large metropolitan centers—exert an asymmetrical gravitational pull on spatial memory. Individuals routinely judge an ordinary or peripheral location as being closer to a salient landmark than that same landmark is to the peripheral location. Such systematic asymmetries render a unified Euclidean coordinate mapping mathematically impossible, as a single coordinate grid cannot simultaneously support non-symmetric inter-point distances.
Furthermore, human spatial judgments routinely violate the triangle inequality, which mandates that the direct distance between two points must be less than or equal to the sum of the distances traversing an intermediate point ($d(A, C) le d(A, B) + d(B, C)$). When physical environments are segmented by political, natural, or categorical boundaries, distances that cross these boundaries are perceived as significantly longer than equivalent distances contained entirely within a single cognitive region. The cognitive overhead required for the brain to maintain a continuous, globally calibrated coordinate matrix across large-scale spaces is computationally prohibitive. Rather than dedicating millions of neurons to constantly recalculate global trigonometry, human cognition avoids global metric synthesis entirely, relying instead on localized, task-dependent, and heuristic spatial estimations.
2.2 Conceptualizing the Cognitive Collage
To replace the flawed metaphor of the metric cognitive map, Tversky introduced the concept of the “cognitive collage” (1993). The cognitive collage captures the fragmented, multi-sourced, and non-uniform nature of human spatial knowledge. Rather than a singular, smoothly rendered atlas, spatial representations resemble a scrapbook or thematic collage: an assemblage of heterogeneous fragments of spatial information acquired at different times, via disparate modalities, and through varying frames of reference. These fragments are loosely bound together without a universal scale, global orientation, or singular perspective.
A cognitive collage effortlessly synthesizes fundamentally diverse inputs. An individual’s spatial knowledge of a metropolis typically combines direct egocentric sensorimotor experiences gained through street-level walking, route-based descriptions provided by transit signage or spoken instructions, bird’s-eye cartographic memories gleaned from printed transit maps or navigational apps, and abstract non-spatial associations such as socioeconomic status or neighborhood safety. These distinct sources are not synthesized into an absolute metric coordinate matrix. Instead, they remain semi-autonomous regional sub-spaces juxtaposed against one another. An individual may possess an exceptionally detailed, route-based representation of their commute, a coarse top-down survey representation of the downtown business district, and an isolated conceptual model of a local park, while remaining entirely uncertain about the precise metric vector connecting them across the city.
The operational triumph of the cognitive collage lies in its pragmatic coherence. Despite the complete absence of global metric accuracy, individuals navigate complex environments, plan intricate multi-stop trajectories, and communicate spatial routes with high functional success. The cognitive collage operates through local consistency rather than global uniformity. For almost all human behavioral objectives—such as boarding the correct subway line, turning at an identifiable intersection, or avoiding an obstacle—a loose patchwork of categorical relationships, topological sequences, and localized visual cues is more than sufficient. By shedding the requirement for global metric integration, the human mind achieves maximum navigational efficiency while minimizing the expenditure of limited computational and working memory resources.
2.3 Cognitive Graphs and Topological Networks
Within Tversky’s conceptual framework, the architectural shift from continuous pictorial fields to discrete structural representations is best understood through the mathematical language of graph theory. While metric cognitive maps require continuous Cartesian planes where entities are positioned along real-valued axes ($x, y, z$), the cognitive collage is structurally instantiated as a cognitive graph—a flexible topological network composed of discrete nodes and relational edges. In this network paradigm, physical entities, distinct environmental intersections, and identifiable landmarks are represented as nodes, while navigational routes, physical pathways, and observed spatial relations constitute the connecting edges.
Topological representations prioritize properties that are invariant under continuous deformation, such as connectivity, adjacency, order, and containment, rather than metric invariants such as absolute distance, exact curvature, or fixed angle. Non-metric spatial reasoning operates directly upon these graph structures. For example, to navigate from point A to point C, a traveler does not need to compute the Euclidean straight-line azimuth; they merely need to traverse a path of connected edges through intermediate node B ($A \rightarrow B \rightarrow C$). Human route-finding relies heavily on this topological logic: the primary objective is knowing which segment connects to which intersection, while the precise metric length of the street or the exact degree of the turn angle can be dynamically resolved using immediate perceptual feedback during locomotion.
Cognitive graphs are organized into hierarchical clusterings within working and long-term memory. Sub-networks representing localized spatial clusters (such as rooms within a building, or specific neighborhoods within a city) are bound together into superordinate structural nodes. Spatial search and path planning operate hierarchically across these layers: long-range navigation first computes coarse routes between superordinate clusters before resolving local paths within lower-level sub-graphs. This hierarchical graph architecture explains why spatial inference is non-uniform: cognitive operations within a tightly clustered topological sub-network exhibit high speed and low error, whereas inferences spanning disconnected or weakly linked sub-networks incur pronounced cognitive delays and systematic distortion.
3. Structural Architecture of Spatial Mental Models
3.1 Spatial Entities and Landmark Categorization
The foundational layer of any spatial mental model consists of the segmentation of continuous physical surroundings into discrete spatial entities. The human perceptual apparatus does not process the world as a uniform field of radiant energy; it relies on Gestalt principles of visual grouping—such as proximity, similarity, continuity, and common fate—to carve continuous space into figures and grounds. In Tversky’s model, this perceptual segmentation produces distinct spatial entities: reference objects, paths, and localized regions. The identity, salience, and behavioral utility of these entities dictate their status within the structural hierarchy of the mental model.
Among these segmented entities, landmarks serve as the foundational cognitive anchors of spatial models. A landmark is not merely a prominent physical object; it is an entity endowed with functional, perceptual, and semantic salience. As demonstrated in seminal research by Sadalla, Burroughs, and Staplin (1980), landmarks act as primary reference points that organize the spatial positions of nearby secondary entities. Spatial relationships are fundamentally asymmetric: secondary, less salient targets are localized relative to stable, highly visible reference anchors (e.g., “the coffee cart is next to the cathedral”), whereas the reverse formulation feels cognitively unintuitive and computationally unnatural (e.g., “the cathedral is next to the coffee cart”). Landmarks anchor spatial memory, creating localized coordinate fields within which surrounding entities are positioned.
In addition to structural anchoring, physical landmarks undergo semantic attribution within long-term memory. Landmarks are rarely encoded as raw geometric shapes; they are enriched with cultural significance, personal history, functional affordances, and administrative boundaries. This semantic layering transforms a raw physical object—such as a clock tower, a mountain peak, or a bridge—into an information-dense cognitive node. The convergence of structural distinctiveness and rich semantic association ensures that landmarks serve as the high-priority retrieval keys within mental models, functioning as navigational waypoints, cognitive boundary markers, and anchors for mental triangulation.
3.2 Qualitative Spatial Relations and Spatial Calculi
Once entities and landmarks are segmented within an environment, the human spatial framework encodes the relationships between them using qualitative spatial relations rather than fine-grained coordinate representations. Fine metric estimations—such as determining that an object is precisely 4.7 meters away at an azimuth of 137 degrees—are computationally fragile and decay rapidly in memory. Instead, human cognition relies on categorical spatial primitives that group continuous metric values into stable, semantically meaningful qualitative buckets. These qualitative abstractions form an intuitive spatial calculus that governs human understanding of layout, containment, and proximity.
The most fundamental qualitative primitives are topological, capturing relationships that remain stable regardless of viewer motion or scale transformations. These primitives align closely with formal qualitative spatial reasoning formalisms, such as the Region Connection Calculus (RCC-8). Human mental models intuitively encode whether entities are disconnected, externally connected (touching), overlapping, or contained within an interior boundary. For example, knowing that a key is inside a desk drawer provides immediate, actionable behavioral utility, whereas the exact distance in centimeters between the key and the drawer’s perimeter is functionally irrelevant for retrieval.
Beyond topology, spatial mental models employ projective relations and qualitative distance tiers. Projective relations establish directional vectors based on an observer’s perspective, an intrinsic object axis, or global reference systems (e.g., in front of, behind, to the left of, above, north of). Concurrently, continuous linear distance is quantized into qualitative distance tiers such as adjacent, near, and far. These categorical boundaries are not defined by rigid metric thresholds; rather, they are modulated by environmental scale, functional interaction, and physical obstacles. Two cities separated by fifty miles may be deemed “near” in the context of transcontinental travel, whereas two offices separated by fifty meters on different floors of a secured building may be deemed “far.” Spatial mental models leverage these qualitative abstractions to construct durable relational architectures that withstand memory degradation.
3.3 Dimensionality and Multi-Tiered Representation
A hallmark of Barbara Tversky’s spatial framework is its multi-tiered and dynamically dimensional nature. Depending on the cognitive task, communicative goal, or behavioral context, the human mind abstracts complex, three-dimensional physical entities into lower-dimensional geometric representations. Dimensionality is not an intrinsic, permanent property of an internal mental entity; rather, it is a fluid cognitive affordance that can be scaled up or down on demand.
Tversky identified that spatial entities are systematically processed through zero-dimensional, one-dimensional, two-dimensional, and three-dimensional representations:
- Zero-dimensional (Point-like) Abstraction: Highly complex, volumetric entities are collapsed into zero-dimensional cognitive points when macro-spatial relationships are computed. For example, in planning a cross-country flight, immense metropolitan areas with sprawling architecture and complex topographies are treated as single, dimensionless point-nodes within a flight network graph.
- One-dimensional (Linear) Abstraction: Environments, connections, and movements are abstracted into one-dimensional lines or paths. This linear abstraction dominates route planning, where the focus is restricted entirely to sequential vectors, turn sequences, and travel corridors, while the width, lateral depth, and surrounding volumetric terrain are ignored.
- Two-dimensional (Surface/Planar) Schemas: Spaces are conceptualized as two-dimensional surfaces, regions, or boundaries. This planar perspective characterizes regional categorization (e.g., territorial maps, state boundaries, or room floorplans), prioritizing containment, surface coverage, and lateral spatial boundaries.
- Three-dimensional (Volumetric) Layouts: Full volumetric representations are reserved for local, highly interactive, and immediate spaces where height, depth, and vertical obstruction directly dictate bodily manipulation, obstacle avoidance, and mechanical intervention.
Crucially, the mind shifts dynamically between these dimensional tiers based on behavioral demands. When describing the distance between two cities, a speaker conceptualizes them as zero-dimensional points. Yet, upon arriving at the city perimeter, the mental model expands into a two-dimensional street network; upon entering an apartment building, it further elaborates into a three-dimensional volumetric space. This dynamic dimension switching enables human cognition to conserve representational resources, expanding to higher-dimensional complexity only when immediate action or communicative precision demands it.
4. The Spatial Framework Analysis: Body-Centered Asymmetries
4.1 The Tri-Axial Model of Human Body Space
To establish how the mind organizes and accesses spatial mental models, Barbara Tversky, in collaboration with Nancy Franklin, developed the Spatial Framework Model (Franklin & Tversky, 1990). The framework asserts that when people mentally imagine or occupy an environment, they construct an internal spatial scaffold based on the three canonical axes of the human body: the head-feet (vertical) axis, the front-back axis, and the left-right axis. Rather than treating internal space as an isotropic sphere where all directions are processed with equal cognitive speed, Tversky proved that retrieval of spatial information across these three axes is fundamentally asymmetric, dictated by the physical morphology of the human body and the ecological realities of the physical world.
The tri-axial model is governed by two interacting constraints: structural asymmetry of the human body and asymmetry of the surrounding physical environment. The head-feet axis is maximally asymmetric both anatomically (the head is structurally and functionally distinct from the feet) and environmentally (anchored by the continuous, omnipresent downward force of gravity). The front-back axis possesses pronounced anatomical asymmetry (our perceptual organs—eyes, nose, mouth—and our primary mechanisms for locomotion and manipulation are oriented toward the front, leaving the back blind and vulnerable), but lacks an asymmetric physical force field comparable to gravity. Finally, the left-right axis is anatomically symmetric (the human body exhibits bilateral symmetry, with identical limbs on either side) and environmentally symmetric (the physical world offers no constant lateral force distinguishing left from right).
Through chronometric experiments measuring the latency of directional verifications, Franklin and Tversky established a reliable cognitive hierarchy. When participants are asked to identify which object is located in a specific direction relative to their imagined body, retrieval speeds directly reflect these bodily and environmental asymmetries:
- Head-Feet (Vertical) Axis: Consistently yields the fastest verification times and the lowest error rates due to the dual reinforcement of anatomical asymmetry and the gravitational vector.
- Front-Back Axis: Demonstrates intermediate retrieval latencies; slower than the vertical axis due to the absence of an environmental force like gravity, but significantly faster than the lateral axis due to acute perceptual and locomotor bodily asymmetry.
- Left-Right Axis: Systematically produces the slowest retrieval speeds and the highest error rates, caused by bilateral anatomical symmetry and the absence of an environmental lateral asymmetry, requiring complex internal discrimination.
4.2 Gravity and the Primacy of the Vertical Axis
The profound cognitive dominance of the vertical axis within human spatial mental models highlights the embodied nature of cognition. While human beings can alter their yaw and pitch, the gravitational force vector remains an invariant, unidirectional constraint throughout evolutionary history. Every human motor command, equilibrium adjustment, and spatial calculation is executed against this omnipresent vertical anchor. Consequently, gravity serves as the primary external reference vector around which spatial perception and mental representation are organized.
This environmental invariance manifests as an acute cognitive ease in discriminating upright versus inverted physical orientations. Visual scenes, facial configurations, and architectural spaces that are rotated away from the gravitational upright incur severe perceptual processing costs. In chronometric reaction-time tasks, identifying spatial relationships that are aligned with the vertical vector requires substantially less working memory activation than identifying relationships on horizontal planes. The human vestibular system—specifically the otolith organs (the utricle and saccule)—continuously signals the absolute direction of gravity to the central nervous system, ensuring that the brain’s internal spatial framework maintains constant vertical alignment.
The primacy of the vertical axis extends far beyond physical locomotion, exerting a powerful influence on temporal and semantic cognition across diverse cultures. Across spoken languages, sign languages, and graphic writing systems, the vertical dimension is consistently privileged. Concepts denoting power, virtue, high social status, positivity, and divine hierarchy are mapped onto the “up” vector, while vulnerability, depression, moral degradation, and subordination are mapped onto the “down” vector. This universal metaphorical scaffolding directly stems from the embodied, physical reality of the vertical axis: rising requires physical effort, energy, and postural vitality against gravity, while falling or descending denotes weakness, collapse, or death. Tversky’s framework demonstrates that the structural mechanics of human embodiment continuously calibrate both physical spatial models and higher-order symbolic thought.
4.3 Allocentric versus Egocentric Reference Systems
To successfully interact with an environment, spatial mental models must negotiate between multiple frames of reference. These frames govern how spatial relations are cognitively configured, stored, and retrieved. Tversky’s paradigm delineates three primary reference systems: egocentric, intrinsic, and allocentric (extrinsic) frameworks, each possessing distinct computational properties and behavioral costs.
An egocentric reference framework is centered on the body of the observer. Spatial coordinates are computed relative to the viewer’s current sensorimotor axes: “the cup is to my left,” “the door is behind me.” Egocentric models are immediate, perceptually grounded, and essential for real-time motor execution, reaching, and path avoidance. However, egocentric representations are highly transient; every time an individual rotates, steps forward, or tilts their head, the entire egocentric coordinate array must be dynamically recalculated.
Conversely, an allocentric (or extrinsic) reference framework specifies spatial relationships independently of the observer, anchoring coordinates to global environmental features, cardinal directions, or fixed structural layouts: “the bank is north of the post office,” “the chair is on the south side of the plaza.” Closely related is the intrinsic framework, which defines spatial relations based on the inherent structural axes of external objects: “the ball is in front of the car” (where “front” is determined by the vehicle’s directional headlights, not the observer’s viewing angle). Allocentric models are cognitively durable and viewpoint-invariant, allowing an individual to retain spatial relationships over long temporal horizons regardless of bodily movement.
The cognitive challenge arises when individuals must translate representations between these reference systems. Shifting from an allocentric map representation to an egocentric navigational command imposes measurable mental costs. For instance, when a driver navigates southward while looking at a standard north-up cartographic map, an egocentric “right turn” corresponds to an allocentric “westward” vector, requiring a 180-degree mental coordinate inversion. Tversky’s research proved that this mental translation generates cognitive interference, elevated error rates, and increased reaction times. Spatial mental models minimize these costs by maintaining flexible, hybrid structures that store structural relations allocentrically while dynamically generating egocentric projections tailored to immediate action contexts.
5. Mental Perspective-Taking and Spatial Transformations
5.1 Dynamic Spatial Updating in Navigation
As human beings traverse complex landscapes, their internal spatial mental models must continuously reconcile changing physical positions with external surroundings. This process, known as dynamic spatial updating, requires the continuous tracking of ego-position and heading vector relative to environmental landmarks, even when those landmarks pass outside the current visual field. Rather than regenerating the spatial model anew at every step, the mind executes continuous transformation algorithms that update spatial relationships in real time.
Spatial updating relies on the integration of two distinct information streams: continuous idiothetic (inertial) path integration and visual-allocentric landmark reconciliation. Idiothetic path integration combines proprioceptive feedback, vestibular acceleration signals, and motor efference copies to compute accumulated displacement and heading changes. This non-visual dead reckoning enables a blindfolded person to point toward an initial starting position after executing several turns. However, path integration is notoriously subject to cumulative drift and error over extended trajectories. To prevent the internal spatial model from drifting catastrophically, the cognitive architecture executes periodic “landmark recalibrations,” reconciling the integrated vector against recognizable visual, auditory, or cartographic reference points.
The sensorimotor costs associated with updating non-visible environmental features are asymmetric and cognitively demanding. Research confirms that when an individual turns 90 degrees to their right, spatial mental models rapidly update the locations of objects that enter their direct visual field. However, objects located behind the individual or in adjacent, occluded rooms require higher cognitive effort to update. In complex, non-orthogonal architectural networks—such as serpentine corridors, skewed intersections, or multi-level subterranean transit hubs—path integration rapidly fails, leading to the phenomenon of heading vector divergence. Here, the internal spatial framework decouples from objective external geography, inducing the acute psychological disorientation characteristic of being lost.
5.2 Mental Rotation and Alignment Effects
The study of spatial transformations within cognitive science was revolutionized by the discovery of mental rotation by Shepard and Metzler (1971), who demonstrated that the time required to determine whether two three-dimensional shapes are identical is a direct, linear function of the angular disparity between them. Barbara Tversky extended the investigation of spatial transformations beyond isolated visual objects, exploring how individuals transform their internal spatial frameworks across environmental, cartographic, and narrative contexts.
A central finding in Tversky’s paradigm is the robust presence of alignment effects in spatial memory and cartographic interpretation. Human spatial reasoning exhibits maximum accuracy and speed when the internal reference framework is physically aligned with the external display. If an individual consults a “you-are-here” directory map that is rotated 180 degrees away from the actual physical environment (for example, a map where “up” corresponds to physical South), profound cognitive disorientation ensues. Under such misaligned conditions, individuals must either execute an exhaustive mental rotation of the entire cartographic display to match the surrounding scene, or mentally rotate their own ego-perspective into the coordinate framework of the map. Both operations impose heavy working memory loads, resulting in pronounced response latencies and frequent navigational errors.
Tversky’s experiments revealed a critical distinction between two mental transformation strategies: rotating the mental representation of the array versus shifting one’s personal imagined perspective. Rotating an entire multi-object environment imposes a high cognitive load that scales steeply with the number of entities within the layout. Conversely, shifting one’s imagined personal perspective—imagining oneself walking to the opposite side of a table or entering a room from a different doorway—proves computationally lighter and more natural. This indicates that human spatial mental models are inherently structured to accommodate ego-motion rather than the holistic rotation of external physical worlds, reinforcing the foundational principle that spatial cognition is evolved to serve bodily movement through stable environments.
5.3 Observer Perspective Variations: Bird’s-Eye versus Field View
When mentally conceiving or recounting an environment, human cognition primarily alternates between two distinct perspectives: the survey perspective (bird’s-eye view) and the route perspective (field or first-person view). These perspectives are not merely stylistic narrative choices; they represent structurally distinct methods of spatial encoding that emphasize fundamentally different types of spatial information.
The survey perspective adopts an external, exocentric, top-down viewpoint looking down upon the terrain from an elevated altitude. This perspective mirrors a cartographic map: it relies on global, two-dimensional coordinate frameworks, cardinal directions, and allocentric spatial relationships (e.g., “The library is north of the central fountain, and the gymnasium is to the west”). The survey perspective provides a comprehensive, structural overview of an environment, explicitly representing distances, boundary contours, and triangular relationships between disconnected locations, facilitating the mental calculation of novel shortcuts.
Conversely, the route perspective adopts an internal, egocentric, sequential viewpoint grounded at the eye-level of a moving observer within the physical landscape. It relies on personal directional terms, landmark sequences, and temporal progression (e.g., “Walk straight down the main path until you see the fountain, then take a sharp right, and the library will be directly ahead”). The route perspective is inherently dynamic, action-oriented, and procedural, prioritizing immediate turn choices, visual vistas, and linear connectivity over global spatial geometry.
Tversky’s empirical investigations into discourse and spatial problem-solving revealed that individuals exhibit dynamic flexibility between these two modes. While the modality of initial spatial acquisition (e.g., learning a layout via a printed street map versus walking through it with a blindfold) biases an individual’s initial perspective choice, human spatial mental models readily translate between route and survey representations. When asked to communicate environmental knowledge, speakers frequently blend both frameworks, utilizing the survey perspective to establish macro-regions and switching to the route perspective to guide localized movement. This dual-perspective capacity demonstrates that underlying spatial mental models are sufficiently abstract to generate either first-person or bird’s-eye representations according to context.
6. Language, Narrative, and Spatial Comprehension
6.1 Linguistic Construction of Spatial Situation Models
One of the most powerful aspects of Tversky’s framework is its capacity to explain how abstract linguistic discourse evokes vivid, functionally active spatial mental models. Human language is inherently linear, acoustic, and symbolic; yet, when reading a narrative or listening to a story, listeners do not retain an inventory of grammatical syntax. Instead, they rapidly translate textual input into a rich, non-linguistic mental representation known as a situation model (van Dijk & Kintsch, 1983). Tversky, working alongside Nancy Franklin, applied the Spatial Framework Model to text comprehension, revealing that narrative descriptions of space trigger the exact same body-centered cognitive scaffolds as direct perceptual experiences.
In classical experiments by Franklin and Tversky (1990), participants were exposed to descriptive narratives placing them within an imagined environment surrounded by various objects (e.g., an opera house with a chandelier above, a rug below, an orchestra pit in front, a balcony behind, a harp to the left, and a velvet curtain to the right). The text then instructed the reader to turn in place to face a different object. Crucially, when subsequently probed with directional verification questions (“What is to your right? What is behind you?”), participants’ reaction times matched the exact tri-axial hierarchy of physical body space: head-feet was verified fastest, followed by front-back, with left-right being the slowest. This established that language comprehenders spontaneously construct a quasi-spatial, body-centered mental model and continuously update it as the narrative unfolds.
This dynamic updating process is mediated by linguistic connectives and narrative cues. When a text states that a protagonist “traversed the hall and entered the laboratory,” readers immediately shift the locus of attention within their internal spatial mental model. Entities located in the vacated room exhibit reduced cognitive accessibility and slower priming latencies, while objects present within the newly occupied room become foregrounded in working memory. Furthermore, spatial proximity within the situation model interacts directly with conceptual processing: objects described as being near the protagonist are retrieved significantly faster than objects located further away, proving that spatial distance within a textually evoked mental model exerts functional constraints on memory accessibility.
6.2 Semantics of Spatial Prepositions and Deixis
Linguistic communication of spatial relationships relies heavily upon spatial prepositions and deictic markers. While formal linguistic semantics initially attempted to define prepositions such as in, on, under, and behind through rigid geometric constraints (e.g., in denoting geometric inclusion within a three-dimensional volume), Barbara Tversky demonstrated that the cognitive semantics of spatial prepositions is deeply functional, contextual, and action-oriented.
Tversky observed that the use of a spatial preposition does not merely report geometric coordinates, but communicates dynamic functional relationships such as support, containment, and control. For example, a pear sitting in an overflowing bowl is described as being “in the bowl” even if it rests geometrically above the physical rim of the vessel, because the bowl structurally supports and constrains the pear. Conversely, an apple placed on top of an inverted bowl is described as “on the bowl.” The semantic choice between in and on is governed by the perceived functional affordance of containment versus surface support, not by a mathematical coordinate intersection. Spatial prepositions capture qualitative physical dynamics, reflecting intuitive human physics and interaction capabilities.
Furthermore, spatial deixis introduces profound perspectives that demand ambiguity resolution. Deictic expressions (such as here, there, this, that, left, and right) are fundamentally indexical; their referential meaning depends entirely upon the position and orientation of the communicative anchor. When a speaker instructs an interlocutor to “look at the statue behind the column,” the phrase contains inherent projective ambiguity: does “behind” mean from the perspective of the speaker (viewer-centered), from the perspective of the listener, or on the intrinsic far side of the column relative to an absolute path of travel? Tversky’s research illustrates how interlocutors actively track shared common ground and conversational context to resolve these deictic ambiguities, demonstrating that spatial language relies upon a collaborative negotiation of shared spatial frameworks.
6.3 Metaphorical Extension of Spatial Schemas to Abstract Thought
Perhaps the most sweeping implication of Tversky’s paradigm is the realization that spatial mental models provide the foundational scaffolding for non-spatial, abstract human thought. While physical navigation across geographic terrain was the evolutionary impetus for developing spatial cognition, the human brain repurposed these spatial computational circuits to navigate abstract conceptual domains. This phenomenon—formalized in George Lakoff and Mark Johnson’s Conceptual Metaphor Theory and deeply expanded by Tversky’s empirical scholarship—is known as the spatialization of thought.
Human beings routinely project abstract, intangible concepts onto spatial schemas. The domain of time is an archetypal example: across nearly all civilizations, temporal duration and progression are spatialized along horizontal and vertical axes. In Western cultures, time is conceptualized as an arrow moving across a horizontal timeline from left to right, matching the directionality of reading and writing. Events in the past are mentally located to the left or “behind,” while the future is positioned to the right or “ahead.” In other cultures, such as the Aymara of the Andes, the past is spatialized in front of the observer (because it is known and can be mentally “seen”), while the future is behind (because it is unknown and unseen). In East Asian scripts that read vertically, earlier events are frequently mapped to higher vertical positions, and later events to lower ones.
Beyond time, spatial frameworks scaffold social, emotional, and organizational domains:
- Vertical Hierarchies: Power, socioeconomic status, and executive control are spatialized vertically, leading directly to institutional expressions like “corporate ladder,” “upper class,” “top-down management,” and “subordinate status.”
- Emotional Valence: Happiness and health are mapped to upward vectors (“feeling up,” “spirits lifted,” “peak performance”), whereas sickness and sorrow are mapped downward (“feeling down,” “plunging into depression,” “falling ill”).
- Quantitative Magnitude: Numerical quantities are structured along a mental number line where increase corresponds to an upward or rightward trajectory.
In her synthesis of these mappings, Tversky advanced the Spatial Metaphor Hypothesis, asserting that spatial schemas are not merely poetic linguistic flourishes, but constitute the cognitive substrate of reasoning itself. Abstract relations become intelligible because the human mind maps them onto the well-understood, sensorimotor structures of physical space, enabling individuals to inspect, manipulate, and infer abstract conclusions using evolutionary spatial cognitive machinery.
7. Systematic Distortions, Heuristics, and Cognitive Biases
7.1 The Alignment Heuristic in Geographic Cognition
Because the human cognitive architecture avoids the computational burden of maintaining a continuous metric coordinate system, it relies heavily on heuristic shortcuts to organize spatial memories. These heuristics simplify, regularize, and compress geographic configurations. One of the most ubiquitous biases identified by Barbara Tversky is the alignment heuristic: the systematic tendency to mentally shift, align, and organize disparate, non-parallel geographic entities so that they cohere along canonical, rectilinear axes (Tversky, 1981).
In her classic investigations of geographic cognition, Tversky presented participants with blank outlines of continental landmasses and asked them to judge relative latitudes and longitudes, or to place them correctly relative to one another. The results revealed striking, systematic errors that directly contradicted geographic reality. For example, the vast majority of participants judged North America and Europe to be geographically aligned across the Atlantic Ocean along equivalent lines of latitude, placing the United States directly across from continental Europe. In physical reality, Europe lies far to the north: the northernmost regions of the contiguous United States (the Canadian border along the 49th parallel) align roughly with Paris, while Rome sits at the same latitude as Chicago, and London aligns with southern Canada.
The alignment heuristic occurs because people represent large-scale geographic landmasses not by their fine-grained coastal contours or true geodesic coordinates, but as superordinate, bounded perceptual units. When multiple bounded units are remembered together, the cognitive system executes an automatic alignment transformation, shifting their relative centroids until they align along a shared horizontal or vertical axis. This cognitive pressure toward perceptual regularity and rectilinear organization reduces the informational entropy of memory storage, but introduces profound, predictable distortions into macro-scale geographic reasoning.
7.2 The Rotation Heuristic and Canonical Orientations
Working in tandem with the alignment heuristic is the rotation heuristic. Tversky demonstrated that when a geographic feature, shoreline, street grid, or continental landmass is tilted at an oblique angle relative to the cardinal environmental axes (North-South, East-West), people systematically mentally rotate the feature toward the nearest canonical vertical or horizontal axis (Tversky, 1981).
A classic empirical demonstration of the rotation heuristic involves the cognitive distortion of the coastline of California. In geographic reality, the California coast curves dramatically from northwest to southeast at a roughly 45-degree angle. However, because individuals intuitively conceive of the Pacific coastline as a vertical, north-south western boundary of the United States, their internal spatial mental models systematically rotate the coastline toward the true North-South vertical axis. This localized rotation heuristic produces famous geographic paradoxes: when asked which city is further east, Reno, Nevada, or San Diego, California, most people intuitively answer Reno, assuming that inland Nevada must be east of coastal Southern California. In reality, because the California coast curves so far to the east, San Diego (117.16° W) is significantly farther east than Reno (119.81° W).
A similar rotation effect distorts the mental representation of South America. The South American continent lies almost entirely east of North America, with the longitude line running through Miami, Florida (80.19° W) passing along the western coast of Ecuador and Peru. However, because both continents are categorized under the shared linguistic and conceptual label of “the Americas,” people mentally rotate South America toward the canonical vertical axis, aligning it directly beneath North America. These rotation heuristics demonstrate that internal spatial mental models prioritize orthogonal, rectilinear simplicity over veridical coordinate orientation, trading geometric fidelity for cognitive simplicity.
7.3 Hierarchical Grouping and Boundary Effects
Human memory is fundamentally structured through hierarchical categorization: specific items are nested within superordinate categories, which are in turn subsumed under higher-order classifications. In physical space, continuous geographic terrain is hierarchically divided by administrative borders, state boundaries, mountain ranges, rivers, and cultural zones. Barbara Tversky, alongside researchers like Stevens and Coupe (1978), demonstrated that these hierarchical groupings profoundly warp distance calculations and directional judgments.
When two physical entities belong to the same superordinate regional category, they are remembered as being significantly closer together than two entities separated by an equivalent Euclidean distance that cross a categorical boundary. For instance, in laboratory experiments where participants estimate distances across artificial city layouts, placing an arbitrary line—designated as a county or state boundary—between two buildings systematically inflates the estimated distance between them. The cognitive boundary functions as a representational barrier that penalizes spatial proximity judgments.
This hierarchical distortion was demonstrated in the classic “Reno-San Diego” style geographic experiments. Stevens and Coupe demonstrated that when participants were asked to judge the directional relationship between San Diego, California, and Reno, Nevada, or between Montreal, Canada, and Seattle, Washington, their judgments were derived from the superordinate categorical relationship between the states or nations rather than the coordinates of the individual cities. Because California is universally understood to be west of Nevada, individuals erroneously infer that any city within California must be west of any city within Nevada. The structural properties of the parent category dictate the spatial attributes of its nested child nodes, completely overriding localized metric facts.
7.4 Landmark Asymmetry and Anchor Effects
The assumption of Euclidean metric symmetry mandates that the subjective distance from location A to location B must equal the distance from B to A. In a definitive series of studies, Barbara Tversky demonstrated that human spatial mental models systematically violate this axiom due to the cognitive asymmetry of landmarks versus ordinary locations. Landmarks serve as foundational reference anchors, bending and warping the subjective metric fabric around them.
In empirical experiments, participants were asked to evaluate the distance between a prominent, highly familiar landmark (e.g., the Eiffel Tower in Paris, or the Empire State Building in New York City) and a nearby, ordinary, non-salient building (e.g., a local neighborhood bakery or corner café). The findings revealed a systematic asymmetry: participants judged the ordinary, non-landmark location as being significantly closer to the prominent landmark than the landmark was to the ordinary location ($d(text{Ordinary}, text{Landmark}) < d(text{Landmark}, text{Ordinary})$). The landmark acts as a powerful cognitive anchor: because it possesses rich, easily retrievable semantic and spatial coordinates, other entities are drawn into its cognitive orbit.
This asymmetric distance judgment directly confirms that spatial mental models are not metric coordinate matrices, but schema-driven relational structures. The prominent anchor serves as the origin point for localized spatial judgments. When evaluating the distance from an ordinary location to a landmark, the cognitive system rapidly accesses the landmark’s rich representation and projects the unknown point onto it. Conversely, when starting from the massive landmark, the ordinary point is vague and poorly resolved in memory, resulting in an inflated distance estimate. This directional asymmetry provides proof that human spatial memory is fundamentally non-metric and governed by cognitive reference points.
8. Embodied Cognition: Mind in Motion and Action Spaces
8.1 Action Spaces: The Body, the Reachable, and the Navigable
A central tenet of Barbara Tversky’s mature theoretical work—crystallized in her foundational volume Mind in Motion: How Action Shapes Thought (2019)—is that spatial cognition is deeply rooted in physical action. Human beings do not conceptualize space as an abstract, uniform continuum. Instead, following the neurocognitive and ecological subdivisions of space, Tversky argues that the mind divides the physical world into three functionally discrete action spaces, each governed by different sensory inputs, cognitive mechanisms, and behavioral affordances:
- The Body Space (Personal Space): This space encompasses the internal physical volume and surface of the human body. It is calibrated through proprioception, kinesthesis, the vestibular apparatus, and somatosensory feedback. Body space is non-visual at its core; it is the immediate, somatic locus of sensation, postural equilibrium, and self-awareness.
- The Reaching Space (Peripersonal Space): Defined as the immediate physical volume surrounding the body that is accessible to the outstretched arms and hands. Reaching space is the domain of tool manipulation, tactile exploration, and direct physical agency. It is governed by a specialized parieto-frontal neural network that computes real-time motor coordinates. Crucially, reaching space is dynamic: when a person holds a tool, such as a hammer or a cane, the mental model of peripersonal space expands to incorporate the operational reach of that instrument.
- The Navigable Space (Extrapersonal / Environmental Space): This macro-scale space extends beyond the immediate reach of the body and cannot be apprehended from a single vantage point. Navigable space must be explored sequentially across time through locomotion, requiring the integration of successive vistas into a unified cognitive collage. It is the space of path integration, map reading, geographic orientation, and urban navigation.
Each of these action spaces possesses distinct psychological properties. Errors that occur when estimating distances within reaching space (where motor feedback rapidly corrects errors) are fundamentally different from distortions that occur across navigable environmental spaces. By demonstrating that the cognitive architecture treats these action spaces with distinct representational calculi, Tversky bridged the gap between low-level motor neurobiology and macro-level spatial cognition.
8.2 Gesture as an Externalized Spatial Mental Model
Within Tversky’s embodied framework, the physical body is not merely an instrument that executes commands generated by an insulated, disembodied brain; the body is an active, external thinking device. The most striking manifestation of this embodiment is spontaneous gesture. When individuals speak, solve spatial puzzles, or mentally plan routes, their hands move spontaneously, even when speaking on the telephone or conversing in the dark where gestures cannot be seen by an interlocutor.
Tversky’s research demonstrates that spontaneous gesturing is not merely a communicative flourish designed to assist the listener; it is a vital mechanism for cognitive offloading and spatial reasoning that directly assists the speaker. Gestures serve as externalized, dynamic spatial mental models. When an individual explains how to navigate an intricate street layout or describes the mechanical operation of an internal combustion engine, their hands spontaneously construct three-dimensional, transient visuospatial inscriptions in the air. These hand movements schematize the problem, explicitly tracing trajectories, establishing reference nodes, representing spatial transformations, and maintaining spatial relations across working memory.
The functional necessity of gesture has been confirmed through rigorous experimental paradigms. When researchers experimentally immobilize participants’ hands—requiring them to solve complex spatial reasoning tasks (such as mental rotation problems, mechanical inference tests, or spatial route explanations) while holding their arms completely still—their performance degrades significantly. Response latencies increase, error rates rise, and verbal descriptions become fragmented. By physically moving the hands through space, the brain recruits motor and premotor cortices to share the computational load of spatial working memory, confirming that bodily action actively shapes the internal processes of spatial thought.
8.3 Action Affordances and Functional Spatial Perceptions
Integrating James J. Gibson’s theory of ecological affordances into cognitive psychology, Tversky emphasizes that spatial perception is never an objective, disinterested recording of geometric dimensions. Rather, what we perceive in an environment is directly modulated by what we can functionally do within that environment. Perceptual space is filtered through the lens of human physical capacity, energetic expenditure, and motor intentionality.
This principle is supported by extensive contemporary research in embodied spatial perception, such as studies led by Dennis Proffitt (2006) demonstrating that physical exhaustion, chronic pain, or carrying a heavy backpack causes observers to judge the steepness of a hill as significantly greater and the distance across a field as significantly longer. Tversky integrates these findings into her framework by showing that motor intentionality calibrates the resolution and scale of spatial mental models. A staircase is not mentally represented merely as an inclined plane with periodic rectilinear notches; it is represented as a structure that affords climbing.
This action-oriented tuning generates a continuous, reciprocal loop connecting bodily movement, environmental feedback, and mental schemas. When a person approaches an obstacle, their mental model does not wait for a complete metric calculation; it triggers immediate motor schemas based on affordances. If an opening is narrower than shoulder-width, the mental model immediately registers a blockage requiring rotation of the torso. Environmental geometry is thus continually translated into behavioral possibilities. The structural architecture of spatial mental models is fundamentally designed to facilitate bodily survival, manipulation, and forward progress through the physical world.
9. External Spatial Representations: Diagrams, Maps, and Visual Inscriptions
9.1 Externalizing Thought: Cognitive Mechanics of Visuospatial Displays
While internal spatial mental models are exceptionally versatile, they are inherently constrained by the severe capacity limitations of human working memory. To transcend these biological constraints, human cultures developed one of their most transformative cognitive technologies: external visuospatial displays. From prehistoric cave paintings and topographical clay tablets to modern architectural blueprints, geographic maps, and interactive data visualizations, human beings externalize their thoughts by inscribing spatial marks onto physical surfaces (Tversky, 2011).
Tversky’s analysis reveals a complementary, isomorphic relationship between internal spatial mental models and external visual artifacts. Just as the brain organizes internal models through qualitative categories, reference anchors, and topological connections, external diagrams eliminate metric noise and elevate structural relations. Visuospatial displays offload immense cognitive burdens: they take complex relational structures that would rapidly overwhelm internal working memory and freeze them onto an enduring, inspectable physical medium. Once externalized on paper or a screen, human perceptual mechanisms—such as rapid visual pattern recognition and spatial scanning—can be brought to bear on complex problems, transforming difficult mental calculations into rapid visual inspections.
To govern the efficacy of external displays, Tversky articulated the Congruence Principle and the Apprehension Principle. The Congruence Principle dictates that the structure and format of an external visualization must correspond naturally to the internal structure of the conceptual domain it represents. For instance, a diagram representing a temporal process must use directional lines or spatial sequences that mirror the mental flow of time. The Apprehension Principle mandates that the visual elements and relations used in a display must be readily and accurately perceived by human perceptual systems. Effective external representations undergo visual pruning: they deliberately strip away extraneous visual realism, photorealistic textures, and exact metric scales in order to emphasize the essential structural elements, ensuring that the visual artifact directly supports human conceptual models.
9.2 The Semiotics and Syntax of Graphic Devices
In her extensive investigations into visual language, Barbara Tversky demonstrated that graphic displays are not arbitrary collections of marks, but are governed by an intuitive semiotics and visual syntax rooted directly in human spatial cognition. People without formal training in graphic design or geometry systematically assign consistent meanings to basic geometric primitives:
- Dots and Points: Spontaneously interpreted as discrete, bounded entities, single events, or localized reference nodes. A dot represents a “thing” stripped of its internal complexity.
- Lines and Paths: Universally interpreted as connections, routes, boundaries, or relationships between entities. A solid line denotes a permanent physical link, adjacency, or structural continuity, while an intersecting line indicates an interaction.
- Arrows and Vectors: Serve as powerful asymmetric glyphs that encode directionality, dynamic motion, temporal succession, causal power, or behavioral flow. An arrow breaks visual symmetry, instructing the viewer’s perceptual system to move in a designated directional vector.
- Enclosures and Boxes: Leverage the topological primitive of interiority to denote membership, category grouping, protection, or systemic containment. Things placed within a common boundary are immediately processed as belonging to the same functional set.
Furthermore, external diagrams exploit foundational Gestalt grouping principles: spatial proximity denotes relational strength, vertical alignment signifies equivalence or categorical hierarchy, and lateral arrangement denotes sequential ordering. Tversky demonstrated that across evolutionary history, diverse and isolated civilizations developed visual charts that spontaneously converge upon these identical spatial-graphic conventions. The syntax of graphic devices is an external projection of the universal spatial framework that human minds use to segment and structure reality.
9.3 Sketching as an Iterative Dialogue with the Mind
Beyond finished maps and polished technical diagrams, Barbara Tversky conducted groundbreaking research into the cognitive mechanics of freehand, iterative sketching. Studying the spontaneous behaviors of architects, structural engineers, artists, and scientific researchers, Tversky revealed that sketching is not merely the passive recording of a pre-formed, pristine internal thought. Rather, sketching constitutes an active, iterative dialogue between the eye, the hand, and the mind (Suwa & Tversky, 1997).
This process operates as a dynamic constructive cycle:
- Depiction: The hand produces a rapid, rough spatial inscription on paper, translating an incomplete, vague internal mental model into an external mark.
- Perceptual Inspection: The eye gazes upon the drawn marks, allowing the visual cortex to parse the physical sketch as an external stimulus.
- Reinterpretation and Creative Discovery: Because a freehand sketch is inherently rough, ambiguous, and imprecise, it inadvertently creates visual juxtapositions, unintended intersections, and emergent spatial groupings that were not present in the original mental concept. The designer “sees” new structural configurations and functional possibilities that the limited capacity of internal working memory could never have generated on its own.
- Revision: The designer modifies the sketch or draws a new iteration, driving the cognitive discovery process forward.
Sketches serve as ambiguous thinking platforms. When architects design complex structural spaces, their internal mental models rapidly reach cognitive saturation due to the limits of spatial working memory. By externalizing partial thoughts onto paper, they clear internal cognitive bandwidth. The external paper becomes an extended working memory buffer, allowing the creator to inspect unintended spatial consequences, test relational alternatives, and foster creative insights that would remain entirely inaccessible within internal cognition alone.
10. Methodological Approaches to Investigating Spatial Mental Models
10.1 Psychometric and Behavioral Reaction-Time Paradigms
To substantiate that internal spatial mental models operate according to distinct structural axes and qualitative relations, Barbara Tversky pioneered rigorous psychometric and chronometric methodologies. Chief among these is the directional verification reaction-time paradigm. In these experiments, participants are placed within an immersive spatial setting—either physical, virtual, or verbally described—and oriented toward a specific heading. Experimenters present auditory or visual probes naming a directional axis (e.g., “above,” “behind,” “left”) and measure the precise latency in milliseconds required for the participant to verify which object is located in that direction.
By conducting thousands of chronometric trials while systematically manipulating body posture (standing upright, lying supine, or reclining on one’s side), Tversky and her colleagues isolated the separate causal contributions of bodily anatomy and environmental gravity. For example, when participants lie supine, the head-feet axis remains aligned with the body’s structural anatomy, but is no longer aligned with the gravitational vector, causing a predictable shift in retrieval speeds. This rigorous chronometric approach provided empirical proof that mental models do not rely on isotropic coordinate searches, but operate through systematically weighted anatomical frameworks.
To complement directional latency experiments, spatial cognitive researchers deploy dual-task interference designs. Grounded in Alan Baddeley’s multi-component working memory model, these paradigms require participants to maintain a spatial mental model while simultaneously performing a secondary visual task (e.g., color discrimination) or a secondary spatial-motor task (e.g., blind motor tracking or spatial tapping). The finding that spatial-motor tasks severely impair spatial mental model maintenance, while purely visual-perceptual tasks cause minimal interference, decisively established that spatial mental models rely upon distinct visuospatial and motoric cognitive subsystems rather than generic visual imagery.
10.2 Sketch-Mapping and Reconstruction Methodologies
To capture the geometric distortions, rotations, and hierarchical clusterings embedded within long-term spatial memory, Tversky popularized the systematic use of sketch-mapping paradigms. In these protocols, participants are provided with blank paper, standardized bounding frames, or isolated landmark reference points, and are instructed to draw geographic environments, room layouts, or recently explored terrains. Rather than analyzing these drawings as mere subjective illustrations, researchers subject them to rigorous quantitative and geometric analysis.
To mathematically quantify the systematic distortions present within hand-drawn cognitive sketches, spatial researchers deploy bidimensional regression analysis (Tobler, 1977; Friedman & Kohler, 2003). Bidimensional regression compares the two-dimensional Cartesian coordinates ($x_i, y_i$) of landmarks depicted in a participant’s sketch map against their true, veridical physical coordinates ($X_i, Y_i$). This computational method extracts precise, standardized parameter coefficients that isolate the specific mathematical transformations occurring in the subject’s mind:
- Scale Factor ($s$): Quantifies the uniform compression or expansion of the internal model relative to real-world space.
- Rotation Angle ($\theta$): Measures the global angular degree to which the participant’s internal representation has rotated away from true cardinal orientation toward a canonical axis.
- Translation Shifts ($a_1, a_2$): Identifies the lateral and vertical displacement of the coordinate origin.
- Distortion Index ($r^2$): Measures the overall configurational fidelity; low values indicate extreme non-Euclidean, topological warping where inter-point distances and local angles have broken down completely.
While sketch-mapping provides rich data regarding the structural contents of memory, researchers must carefully control for confounding factors such as manual drawing proficiency, hand tremors, and motor execution constraints. To address these motor skill limitations, modern paradigms utilize interactive virtual layout reconstruction tasks. Participants position 3D environmental entities within standardized virtual sandboxes using a digital controller, allowing researchers to measure coordinate placements, rotational alignments, and boundary-induced clustering errors with millimeter precision without requiring drawing ability.
10.3 Verbal Protocols and Linguistic Elicitation
Recognizing the profound convergence between language and spatial representation, Tversky established linguistic elicitation as an indispensable methodological window into the human mind. Rather than relying solely on non-verbal motor tasks, researchers analyze the spontaneous structure, sequence, and vocabulary of spoken narratives when participants are asked to describe environments, give navigational directions, or construct visual scenes.
One of Tversky’s foundational discoveries via this methodology was the identification of universal structural grammars in environmental descriptions. When asked to spontaneously describe an apartment, a town, or a university campus, people do not output a random, disorganized list of items. Instead, they structure their verbal protocols around two highly systematic organizational frameworks: tours (route-based descriptions, moving sequentially through doors, down hallways, and around intersections) or maps (survey-based descriptions, partitioning space into macro-regions and employing global spatial coordinates). The participant’s spontaneous choice of perspective reveals how their internal cognitive collage has been organized, providing critical insights into how the modality of initial learning shapes long-term spatial mental models.
In modern paradigms, linguistic elicitation is paired with high-precision mobile eye-tracking. By tracking gaze fixations while participants speak, researchers can correlate visual attention directly with spoken output. When an individual verbally articulates an abstract relationship (e.g., “The economic hierarchy collapsed”), their eye movements spontaneously trace downward vertical paths, mirroring the embodied spatial schema underlying the linguistic metaphor. These multi-modal protocols bridge the gap between spoken semantics, visual-motor scanning, and internal spatial architecture.
11. Interdisciplinary Applications of the Spatial Mental Models Framework
11.1 Human-Computer Interaction, Interface Design, and UX
The principles articulated in Barbara Tversky’s spatial mental models framework have exerted a profound, transformative influence on Human-Computer Interaction (HCI) and user interface (UX) design. Modern digital computing owes its intuitive ubiquity to spatial metaphors: the foundational “desktop” metaphor—pioneered by Xerox PARC and popularized by Apple and Microsoft—relies entirely on mapping abstract file storage systems onto the peripersonal reaching space of an office desk. Users manipulate “files,” drop them into “folders,” and discard items into a “trash can,” translating complex data operations into intuitive physical interactions.
By leveraging Tversky’s principles of qualitative spatial calculi and structural congruence, interface designers build digital environments that align with human spatial cognition. Interfaces that respect the Congruence Principle—ensuring that the visual arrangement of digital controls directly reflects their functional dependencies—yield significantly lower error rates and reduced cognitive load. For instance, in data dashboards and complex software suites, designers utilize containment boxes to signal functional grouping, visual arrows to guide operational workflows, and vertical hierarchies to communicate systemic permissions.
In the contemporary era of spatial computing—encompassing Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR)—Tversky’s work is essential. VR systems that force users to navigate digital environments using continuous, high-speed linear translations frequently induce acute cyber-sickness and cognitive disorientation caused by vestibular-visual heading vector divergence. By applying the cognitive collage framework, XR architects design locomotion systems around discrete “teleportation” to salient visual anchors, topological wayfinding, and aligned “you-are-here” personal radar maps, minimizing sensorimotor conflict and enabling seamless navigation across infinite virtual terrains.
11.2 Architecture, Urban Design, and Wayfinding Systems
In the domain of urban design and architectural planning, Tversky’s spatial framework provides the theoretical foundation for creating legible, navigable built environments. Her work directly intersects with and expands upon urban theorist Kevin Lynch’s seminal framework in The Image of the City (1960), which categorized urban environmental elements into paths, edges, districts, nodes, and landmarks. Tversky provided the rigorous psychological and experimental validation for Lynch’s observational insights, proving that these five elements constitute the foundational building blocks of the human cognitive collage.
When architects design large-scale, complex civic structures—such as international airports, multi-wing medical centers, or subterranean transit hubs—ignoring human spatial mental models leads to severe wayfinding dysfunction. If an architectural layout relies on non-orthogonal, winding corridors that lack visual access to external vertical reference points (like the sky or the horizon), human idiothetic path integration rapidly breaks down, leaving visitors profoundly disoriented. Incorporating Tversky’s findings, architectural wayfinding strategists actively introduce distinct, highly salient landmarks at primary transit junctions, construct clear visual axes that visually penetrate multiple building levels, and place schematic maps at points of critical decision-making.
A classic urban embodiment of Tversky’s principles is Harry Beck’s famous 1931 redesign of the London Underground map. Prior to Beck, transit maps were drawn to strict geographic, Euclidean scale, resulting in cluttered central clusters and absurdly extended peripheral lines that were almost unreadable. Beck recognized that a subway passenger has zero need for true Euclidean distances or exact surface curvature; they require only a topological cognitive graph. Beck stripped away geographic scale, regularized all lines to 45- and 90-degree angles, enlarged the dense central interchange network, and preserved only topological connectivity and station order. Beck’s schematic map was initially rejected by transit executives as “too radical,” but became an immediate global triumph among passengers because it mirrored the human mind’s internal spatial mental model.
11.3 Artificial Intelligence, Robotics, and Computational Spatial Reasoning
Within computational science and artificial intelligence, Tversky’s paradigm has driven a paradigm shift in how autonomous agents represent, navigate, and reason about physical space. Historically, autonomous mobile robotics relied strictly on dense, metrically precise Simultaneous Localization and Mapping (SLAM) algorithms. These systems construct massive, millimeter-accurate 3D point-cloud matrices or continuous occupancy grids. However, these pure metric systems are computationally brittle: minor sensor calibration errors compound over time, loop-closure failures cause catastrophic coordinate divergence, and the computational storage required scales unsustainably in expansive environments.
To overcome these limitations, roboticists turned to Qualitative Spatial Representation and Reasoning (QSRR) and hybrid semantic topological mapping, heavily inspired by Tversky’s cognitive collage theory. Instead of relying exclusively on global metric coordinates, modern autonomous systems construct multi-tiered hierarchical models:
- Metric Layer: A highly localized, temporary metric occupancy grid reserved strictly for immediate obstacle avoidance, wheel odometry, and reactive path execution within reaching/peripersonal space.
- Topological Layer: A persistent cognitive graph that abstracts rooms, corridors, and intersections into discrete nodes and traversable edges, enabling ultra-fast graph-search path planning across large facilities.
- Semantic Layer: An attribute network that tags topological nodes with functional definitions, object affordances, and human-readable landmarks (e.g., “kitchen,” “charging dock,” “reception desk”).
Furthermore, in the burgeoning field of Vision-and-Language Navigation (VLN) and Embodied AI, artificial agents must comprehend natural language instructions provided by humans (such as “Go past the cafeteria, turn left at the bronze statue, and wait in the lobby”). Robotic agents running purely on Cartesian coordinates fail catastrophically at this task. By incorporating Tversky’s spatial mental models framework, researchers train deep reinforcement learning architectures to parse qualitative prepositions, identify deictic reference anchors, and translate linear linguistic narratives into dynamic internal topological networks, bridging the gap between human spatial discourse and autonomous machine action.
11.4 Pedagogy and STEM Education
The realization that spatial mental models scaffold abstract thought has transformed pedagogical theory, particularly within Science, Technology, Engineering, and Mathematics (STEM) education. Foundational meta-analyses, such as Uttal et al. (2013), demonstrate that spatial reasoning ability is a powerful, reliable predictor of long-term academic success and professional attainment in STEM disciplines. Crucially, spatial ability is not an immutable, genetically fixed trait; it is a highly malleable cognitive skill that can be developed through deliberate instruction.
Applying Barbara Tversky’s research on graphic devices and sketching, educational researchers develop visual literacy curricula that teach students how to read, construct, and critique scientific visualizations. In chemistry, understanding molecular geometry requires students to mentally translate flat, two-dimensional Lewis dot diagrams into dynamic, three-dimensional stereochemical models. In geology and civil engineering, students must mentally slice through two-dimensional topographic contour maps to visualize internal structural fault lines and subsurface fluid dynamics. Instructional strategies that leverage Tversky’s Congruence Principle—ensuring that diagrammatic conventions visually match physical causal mechanisms—significantly improve conceptual comprehension and problem-solving transfer.
Furthermore, pedagogical interventions actively incorporate embodied action and gesture training into the classroom. When educators actively use spatialized gestures to explain abstract mathematical transformations—such as tracing parabolic trajectories through the air or modeling coordinate rotations with their hands—students grasp the underlying concepts more rapidly. Similarly, encouraging students to sketch external representations iteratively while solving complex physics problems prompts the same reflective, creative dialogue observed among professional architects, allowing them to offload working memory, identify hidden assumptions, and master abstract scientific reasoning.
12. Critical Appraisals, Contemporary Extensions, and Future Directions
12.1 Neurocognitive Foundations and Neural Correlates
While Barbara Tversky’s spatial mental models framework was formulated primarily through behavioral, psycholinguistic, and cognitive experimentation, modern neuroscience has provided remarkable neurobiological validation for her functionalist, non-Euclidean paradigm. The biological machinery underlying human spatial cognition resides within an intricately interconnected network encompassing the hippocampal formation, the entorhinal cortex, the retrosplenial cortex, and the posterior parietal lobule.
The discovery of place cells in the hippocampus by John O’Keefe (1971) and grid cells in the medial entorhinal cortex by Edvard and May-Britt Moser (2005) provided the neural substrate for spatial representation. Initially, grid cells—which fire in remarkably regular, triangular tessellations across physical space—were hailed by classical theorists as the long-sought, internal Euclidean metric map. However, subsequent neurophysiological research has decisively vindicated Tversky’s cognitive collage model. Studies have demonstrated that in non-rectangular or complex, fragmented environments, grid cell firing fields warp, stretch, shear, and fragment across environmental boundaries. Rather than maintaining an inflexible, universal metric coordinate system, the entorhinal grid network recalibrates locally around prominent geometric boundaries and salient landmarks, operating precisely as a set of localized, topological sub-maps.
Furthermore, functional neuroimaging (fMRI) studies have demonstrated how the brain negotiates reference frame transformations. The posterior parietal cortex processes egocentric visual and sensorimotor coordinates, while the hippocampus and parahippocampal place area (PPA) maintain allocentric representations. The retrosplenial cortex (RSC) functions as the critical translation hub, converting egocentric perspectives into allocentric models and vice versa. Most profoundly, groundbreaking neuroimaging work by Constantinescu, O’Reilly, and Behrens (2016) proved that human grid-cell and place-cell networks fire not only when navigating physical space, but when navigating abstract, non-spatial conceptual spaces (such as navigating a continuous two-dimensional parameter space of arbitrary bird features). The brain’s navigational circuits serve as the universal engine for general conceptual thought, providing profound biological proof for Tversky’s Spatial Metaphor Hypothesis.
12.2 Theoretical Critiques and Alternative Paradigms
Despite its vast explanatory power, Barbara Tversky’s spatial mental models framework has been the subject of ongoing theoretical debate and critical interrogation within cognitive science. The primary point of contention revolves around the degree of metric preservation within long-term spatial memory. Proponents of continuous vector representations—such as William H. Warren and colleagues—argue that while human spatial judgments exhibit heuristic distortions and localized errors, the brain nonetheless maintains an underlying, continuous vector-based representation that enables smooth, accurate shortcutting across open fields where categorical boundaries are absent. These critics contend that Tversky’s experiments, which frequently utilized artificial textual scenarios or hand-drawn sketches, may systematically overestimate the fragmented nature of human spatial memory by introducing communicative and task-specific demands.
Another theoretical critique addresses the conceptual boundaries separating spatial mental models from generalized schema theory and traditional mental imagery. Some computational modelers argue that Tversky’s definition of the “cognitive collage” is overly flexible and descriptive, functioning more as a compelling post-hoc metaphor than a predictive, mathematically formal algorithmic model. Connectionist and neural network theorists have demonstrated that many of the heuristic distortions documented by Tversky—such as the alignment and rotation heuristics—can naturally emerge from simple, distributed associative learning networks without requiring explicit, rule-based qualitative calculi or multi-tiered categorical abstractions.
Additionally, ecological psychologists working in the strict tradition of J.J. Gibson question the necessity of positing rich internal mental representations at all. Radical embodied cognitive scientists argue that navigational wayfinding does not rely on consulting internal cognitive collages, but is achieved through direct, online perceptual-motor coupling with the environment—such as following continuous optic flow gradients, visual affordances, and behavioral beacons. While Tversky consistently incorporated affordances and action spaces into her scholarship, the debate between representational mental modeling and radical non-representational embodiment remains a vibrant frontier in cognitive theory.
12.3 Future Trajectories in Spatial Cognition Research
As human civilization accelerates into an era defined by ubiquitous digital connectivity, artificial intelligence, and extended reality, Barbara Tversky’s spatial mental models framework offers a vital roadmap for future scientific exploration. One of the most urgent contemporary research trajectories concerns the cognitive impact of digital navigation dependencies. With billions of individuals relying exclusively on GPS-enabled turn-by-turn smartphone navigation (e.g., Google Maps, Apple Maps), the human brain is largely relieved of the need to actively perform idiothetic path integration, attend to environmental landmarks, or construct internal cognitive collages. Longitudinal cognitive studies are currently investigating whether chronic GPS reliance induces neuro-anatomical atrophy within the hippocampus and compromises general spatial reasoning abilities, transforming human spatial cognition from active mental modeling to passive, reactive prompt-following.
Another frontier lies in the exploration of spatial cognition within non-Euclidean virtual environments. Immersive virtual technologies allow computer scientists to construct impossible spaces—interiors that are physically larger than their exteriors, portal-linked spaces that violate spatial transitivity, and environments characterized by hyperbolic or spherical geometries. Investigating how human spatial mental models adapt, fracture, or restructure when exposed to environments that systematically violate terrestrial physics will yield unprecedented insights into the fundamental plasticity and boundary conditions of human spatial reasoning.
Finally, contemporary cognitive science is working to synthesize Tversky’s spatial mental models framework with the overarching computational architecture of predictive processing and active inference (Clark, 2016; Friston, 2010). Under this unified computational view, spatial mental models function as top-down generative models that continuously predict the sensorimotor consequences of physical actions. When sensory feedback diverges from the spatial model’s predictions, the brain computes a prediction error, prompting either a rapid update to the internal cognitive collage or an active physical movement to sample new sensory data. Merging Tversky’s qualitative, action-oriented spatial schemas with the formal mathematical machinery of Bayesian active inference promises to establish a unified grand theory of how the mind perceives, navigates, and imagines the infinite spaces of physical and conceptual reality.
Conclusion
Barbara Tversky’s Spatial Mental Models Framework represents one of the most profound and enduring paradigm shifts in modern cognitive science. By dismantling the long-held myth of the unified metric cognitive map, Tversky liberated spatial psychology from the constraints of rigid Euclidean geometry. She revealed that human spatial cognition is not a passive, photorealistic internal mirror of the physical world, but an active, dynamic, and pragmatic construction—a cognitive collage woven from heterogeneous perceptual fragments, bodily asymmetries, categorical primitives, and linguistic narratives.
Through her identification of the Spatial Framework Model, Tversky proved that the internal architecture of thought is deeply embodied, indelibly marked by the physical morphology of our bodies, the downward pull of gravity, and the functional realities of forward locomotion. Her documentation of systematic cognitive distortions—such as the alignment, rotation, and hierarchical grouping heuristics—demonstrated that human spatial memory is fundamentally organized to conserve computational resources, trading absolute metric precision for high operational efficiency. Furthermore, by tracing the trajectory of spatial thought as it externalizes into spontaneous gestures, architectural sketches, transit maps, and abstract linguistic metaphors, Tversky illuminated the foundational truth that spatial reasoning serves as the overarching cognitive scaffolding for abstract human conceptualization.
As contemporary science grapples with the neurobiological basis of navigation, the development of embodied artificial intelligence, the design of immersive virtual worlds, and the cognitive consequences of digital technology, Barbara Tversky’s insights remain exceptionally relevant. Her scholarship teaches us that we do not simply live within space; space lives within us. It is the silent, pervasive architecture of our thoughts, the living language of our hands, and the foundational matrix through which the human mind makes sense of itself and the universe it explores.
References
- Baddeley, A. D. (2007). Working memory, thought, and action. Oxford University Press. https://doi.org/10.1093/acprof:oso/9780198523192.001.0001
- Clark, A. (2016). Surfing uncertainty: Prediction, action, and the embodied mind. Oxford University Press. https://doi.org/10.1093/acprof:oso/9780190217013.001.0001
- Constantinescu, A. O., O’Reilly, J. X., & Behrens, T. E. J. (2016). Organizing conceptual knowledge in humans with a gridlike code. Science, 352(6292), 1464–1468. https://doi.org/10.1126/science.aaf7860
- Franklin, N., & Tversky, B. (1990). Searching imagined environments. Journal of Experimental Psychology: General, 119(1), 63–76. https://doi.org/10.1037/0278-7393.16.1.63
- Friedman, A., & Kohler, B. (2003). Bidimensional regression: Assessing the ground-truth correspondence and accuracy of cognitive maps. Spatial Cognition & Computation, 3(2-3), 127–154. https://doi.org/10.1207/S15427633SCC032&3_03
- Hafting, T., Fyhn, M., Molden, S., Moser, M. B., & Moser, E. I. (2005). Microstructure of a spatial map in the entorhinal cortex. Nature, 436(7052), 801–806. https://doi.org/10.1038/nature03721
- Johnson-Laird, P. N. (1983). Mental models: Towards a cognitive science of language, inference, and consciousness. Harvard University Press. https://psycnet.apa.org/record/1983-28956-000
- Kosslyn, S. M. (1981). The medium and the message in mental imagery: A theory. Psychological Review, 88(1), 46–66. https://doi.org/10.1037/0033-295X.88.1.46
- Lakoff, G., & Johnson, M. (1980). Metaphors we live by. University of Chicago Press. https://press.uchicago.edu/ucp/books/book/chicago/M/bo24408802.html
- Lynch, K. (1960). The image of the city. MIT Press. https://mitpress.mit.edu/9780262620017/the-image-of-the-city/
- O’Keefe, J., & Dostrovsky, J. (1971). The hippocampus as a spatial map: Preliminary evidence from unit activity in the freely-moving rat. Brain Research, 34(1), 171–175. https://doi.org/10.1016/0006-8993(71)90358-1
- Proffitt, D. R. (2006). Embodied perception and the economy of action. Perspectives on Psychological Science, 1(2), 110–122. https://doi.org/10.1111/j.1745-6916.2006.00008.x
- Pylyshyn, Z. W. (1973). What the mind’s eye tells the mind’s brain: A critique of mental imagery. Psychological Bulletin, 80(1), 1–24. https://doi.org/10.1016/0010-0277(73)90022-1
- Randell, D. A., Cui, Z., & Cohn, A. G. (1992). A spatial logic based on regions and connection. Proceedings of the 3rd International Conference on Principles of Knowledge Representation and Reasoning (KR’92), 165–176. https://doi.org/10.1016/0954-1810(92)90039-6
- Sadalla, E. K., Burroughs, W. J., & Staplin, L. J. (1980). Reference points in spatial cognition. Journal of Experimental Psychology: Human Learning and Memory, 6(5), 516–528. https://doi.org/10.1016/0010-0285(80)90018-0
- Shepard, R. N., & Metzler, J. (1971). Mental rotation of three-dimensional objects. Science, 171(3972), 701–703. https://doi.org/10.1126/science.171.3972.701
- Stevens, A., & Coupe, P. (1978). Distortions in judged spatial relations. Cognitive Psychology, 10(4), 422–437. https://doi.org/10.1016/0010-0285(78)90014-4
- Suwa, M., & Tversky, B. (1997). What do architects and students perceive in their design sketches? A protocol analysis. Design Studies, 18(4), 385–403. https://doi.org/10.1016/S0010-0285(02)00511-7
- Tobler, W. R. (1977). Bidimensional regression. Geographical Analysis, 9(1), 1–20. https://doi.org/10.1111/j.1538-4632.1977.tb00573.x
- Tolman, E. C. (1948). Cognitive maps in rats and men. Psychological Review, 55(4), 189–208. https://doi.org/10.1037/h0061626
- Tversky, B. (1981). Distortions in memory for maps. Cognitive Psychology, 13(3), 407–433. https://doi.org/10.1016/0010-0285(81)90005-8
- Tversky, B. (1993). Cognitive maps, cognitive collages, and spatial mental models. In A. U. Frank & I. Campari (Eds.), Spatial Information Theory: A Theoretical Basis for GIS (pp. 14–24). Springer Berlin Heidelberg. https://doi.org/10.1016/0010-0285(93)90011-R
- Tversky, B. (2011). Visualizing thought. Topics in Cognitive Science, 3(3), 499–535. https://doi.org/10.1111/j.1756-8765.2010.01113.x
- Tversky, B. (2019). Mind in motion: How action shapes thought. Basic Books. https://www.basicbooks.com/titles/barbara-tversky/mind-in-motion/9780465093069/
- Uttal, D. H., Meadow, N. G., Tipton, E., Hand, L. L., Alden, A. R., Warren, C., & Newcombe, N. S. (2013). The malleability of spatial skills: A meta-analysis of training studies. Psychological Bulletin, 139(2), 352–402. https://doi.org/10.1037/a0030383
- van Dijk, T. A., & Kintsch, W. (1983). Strategies of discourse comprehension. Academic Press. https://psycnet.apa.org/record/1983-28956-000