During the early decades of the twentieth century, experimental psychology was gripped by a dogmatic mechanization of the animal and human mind. Under the unyielding banner of classical behaviorism, championed by figures such as John B. Watson, internal mental phenomena—desires, mental representations, hypotheses, and goals—were systematically excised from the scientific lexicon as untestable relics of prescientific animism. The prevailing orthodoxy asserted that all behavior could be exhaustively parsed into reflexive chains of peripheral stimulus-response (S-R) connections stamped into the nervous system through mechanical reinforcement and habituation. Learning was treated as an unmediated, physiological etching process: organisms were passive biological automatons propelled strictly by environmental inputs and immediate physiological drives, leaving no epistemological space for deliberate contemplation, spatial awareness, or autonomous purpose.
Emerging from this austere landscape, Edward Chace Tolman mounted one of the most sustained, intellectually rigorous, and ultimately transformative revolts in the history of behavioral science. Working out of his laboratory at the University of California, Berkeley, Tolman rejected both the ungrounded introspectionism of early German structuralism and the radical peripheral reductionism of American behaviorism. Instead, he pioneered an innovative theoretical orientation termed purposive behaviorism. Tolman contended that behavior is fundamentally molar, holistic, and goal-directed. Rather than merely acquiring rigid motor habits through the blind mechanics of trial and error, animals cultivate comprehensive internal knowledge structures of their environments. In this framework, organisms do not simply respond to stimuli; they actively perceive relationships, construct cognitive hypotheses, and build structured, internal representations of their physical worlds.
The conceptual anchors of Tolman’s revolution were the twin discoveries of latent learning and the cognitive map. Through an ingenious series of experimental paradigms utilizing standardized rodent mazes, Tolman and his collaborators empirically proved that organisms can acquire rich, highly detailed information about spatial environments in the complete absence of primary biological reinforcement. This acquired knowledge remains behaviorally latent, awaiting appropriate motivational conditions to be translated into observable performance. Published formally in his monumental 1948 treatise, “Cognitive Maps in Rats and Men,” Tolman synthesized decades of empirical research to demonstrate that organisms navigate through internal mental representations that encode spatial coordinates, environmental contingencies, and vector geometries. This blog post provides an exhaustive, granular exploration of Tolman’s life, his theoretical innovations, the seminal empirical experiments that shattered the S-R hegemony, and his profound, enduring legacy across contemporary neurobiology, cognitive science, and artificial intelligence.
1. Introduction to Edward Chace Tolman and the Purposive Behaviorist Paradigm
1.1 Biographical Context and Intellectual Trajectory
Edward Chace Tolman was born in 1886 in West Newton, Massachusetts, into an intellectual environment characterized by civic reform, academic rigor, and classical nineteenth-century liberal ideals. His brother, Richard Chace Tolman, would achieve international renown as a mathematical physicist and physical chemist at the California Institute of Technology, reflecting the family’s intense commitment to empirical and scientific pursuits. Edward Tolman initially pursued an undergraduate education in electrochemistry at the Massachusetts Institute of Technology, graduating in 1911. However, after reading the philosophical and psychological writings of William James, Tolman underwent a profound intellectual reorientation. Recognizing that his true passions lay in deciphering the mechanisms of mind, consciousness, and behavior, he enrolled in graduate studies at Harvard University’s Department of Philosophy and Psychology.
Tolman’s graduate tenure at Harvard took place during a momentous epoch of intellectual transition. He studied under prominent philosophical figures such as Ralph Barton Perry and Edwin Bissell Holt, both of whom were instrumental in developing “new realism”—a philosophical movement asserting that conscious experience and relational meaning could be understood objectively as direct encounters with external physical realities, rather than as solipsistic mental projections. Holt’s early attempts to conceptualize the purposive qualities of behavior from an objective perspective exerted a foundational influence on young Tolman. Crucially, in the summer of 1912, Tolman traveled to Giessen, Germany, where he studied under Kurt Koffka, one of the primary architects of Gestalt psychology. This early immersion in Gestalt theory exposed Tolman to the principles of perceptual organization, field theory, and relational holistic structures—concepts that stood in stark opposition to the atomistic, elementistic psychology then dominating Anglo-American laboratories.
Returning to Harvard, Tolman completed his doctoral dissertation on retroactive inhibition in 1915 under the supervision of Hugo Münsterberg. Following a brief appointment at Northwestern University—which ended prematurely due to his vocal pacifist stance during the First World War—Tolman accepted a faculty position at the University of California, Berkeley, in 1918. It was at Berkeley that Tolman established his legendary rodent laboratory and formulated his life’s work. Positioned between the radical peripheral behaviorism sweeping the United States and the introspective mentalism fading in continental Europe, Tolman occupied a unique theoretical nexus. He steadfastly refused to abandon the objective, empirical, and quantifiable methods championed by behaviorism, yet he adamantly rejected its mechanistic, anti-cognitive dogmatism. Tolman’s intellectual journey was characterized by this delicate, brilliant synthesis: constructing an experimentally rigorous science of behavior that recognized the undeniable presence of internal cognitive structures, expectations, and intrinsic purpose.
1.2 The Core Philosophy of Purposive Behaviorism
The philosophical centerpiece of Tolman’s theoretical enterprise was purposive behaviorism, a framework formally codified in his 1932 magnum opus, Purposive Behavior in Animals and Men. Classical Watsonian behaviorism had proclaimed that psychology could only attain scientific legitimacy by banishing all teleological language; behavior was strictly to be viewed as a passive, involuntary consequence of preceding environmental triggers. Tolman completely dismantled this assumption. He argued that behavior does not merely consist of fragmented physical twitches elicited by incoming sensory energy; rather, behavior is intrinsically and objectively goal-directed. An organism’s movements through space are organized around targets, destinations, objects to be reached, or hazards to be avoided.
Crucially, Tolman demonstrated that teleology could be incorporated into empirical, objective behavioral science without resorting to the unscientific metaphysics of vitalism or subjective introspection. For Tolman, purpose was not a mystical, metaphysical fluid driving the organism from within, nor was it a private conscious state accessible solely through introspective report. Instead, purpose was an objectively observable, functional characteristic of the behavior itself. When an animal runs through an intricate maze, its behavior displays persistent variation across changing environmental barriers until a specific terminus—the food box—is attained. If a path is blocked, the animal selects an alternative path; if the terrain changes, the motor mechanics change, yet the directional orientation toward the goal remains invariant.
This functional invariance toward a specific environmental end-state is precisely what defines purposive behavior. Purpose, therefore, is an empirical property inherent in the descriptive organization of action. Behavior, as Tolman famously observed, “reeks of purpose and cognition.” By redefining purpose as an objective, descriptive attribute rather than an occult internal ghost, Tolman circumvented the philosophical trap that had led Watson to reject cognitive constructs altogether. Purposive behaviorism asserted that an experimenter could remain entirely committed to rigorous external observation, measurement, and experimental manipulation while simultaneously recognizing that the organism operates as a purposive, cognitive entity navigating an organized psychological field.
1.3 Molar vs. Molecular Behaviorism
To establish the empirical legitimacy of purposive behaviorism, Tolman formulated a profound ontological distinction between molecular behaviorism and molar behaviorism. Molecular behaviorism—championed implicitly by Watson and later formalized mechanistically by Clark L. Hull—sought to reduce all animal activity to its ultimate physiological constituents: microscopic muscle contractions, neurological reflex arcs, and glandular secretions. In the molecular paradigm, an animal pressing a lever or turning left in a maze is conceptualized as a precise, deterministic sequence of efferent motor impulses triggering specific flexor and extensor muscular twitches. The goal of molecular psychology was the comprehensive reduction of behavior to the underlying physics and physiology of the neuromuscular apparatus.
Tolman argued that molecular reductionism was not only practically impossible for complex actions, but epistemologically bankrupt. He defined molar behavior as holistic, descriptive units of activity that possess emergent properties of their own. A rat traversing a maze, a bird constructing a nest, or a human driving an automobile cannot be understood merely as an aggregate of isolated muscular twitches, just as the functional properties of water cannot be understood by examining isolated hydrogen and oxygen atoms in the absence of their emergent molecular bonds. Molar behavior possesses distinctive, identifiable characteristics: it is directed toward an environmental object or state, it exploits environmental features as means to reach ends, and it exhibits behavioral plasticity—the capacity to alter physical motor patterns dynamically if the environment shifts.
The methodological implications of this molar distinction for animal psychology and maze learning experiments were revolutionary. In a molecular framework, if an animal learned to run through a maze by executing a sequence of right and left turns, that learning was thought to be stored as kinesthetic “muscle memory”—a rigid chain of peripheral motor reflexes where the proprioceptive feedback from one muscular contraction served as the stimulus for the next. Tolman exposed the absurdity of this view. If a maze was flooded with water, a rat trained to run the corridors would instantly swim the maze flawlessly, executing entirely novel muscular movements that it had never practiced during training. Because the molar goal (reaching the destination) and the molar knowledge (the layout of the space) remained constant, the organism could deploy completely novel motor repertoires to achieve the same functional outcome. Molar behaviorism emancipated psychological science from the confines of the peripheral nervous system, elevating the analysis of behavior to the level of holistic, environmental adaptation.
2. The Historical Matrix: Behaviorism, Neobehaviorism, and the Anti-Mentalist Hegemony
2.1 The Classical Watsonian Orthodoxy
The historical backdrop against which Tolman launched his theoretical insurgency was defined by the aggressive rise of classical behaviorism. In 1913, John B. Watson published his seminal manifesto, “Psychology as the Behaviorist Views It,” initiating an intellectual purge of the discipline. Watson proclaimed that psychology had failed to become a true natural science due to its historical reliance on consciousness, introspective methods, and unobservable mental states. He demanded the absolute expulsion of concepts such as mind, feeling, imagery, and ideation from the psychological vocabulary. In Watson’s programmatic vision, human and animal organisms were merely complicated physical machines whose actions were exclusively determined by direct environmental inputs operating through conditioned reflexes.
Watsonian orthodoxy was rooted in strict peripheralism. Watson maintained that even higher cognitive processes, such as human thinking, were nothing more than covert motor habits—specifically, subvocal speech involving minute, microscopic movements of the vocal cords and larynx. Complex animal navigation was interpreted purely as a concatenated chain of peripheral motor habits. In maze learning, a rat was assumed to acquire a mechanical succession of kinesthetic associations: the kinesthetic sensations generated by running down the first alley mechanically triggered the motor command to execute a ninety-degree turn at the intersection. Learning was conceptualized as a strictly peripheral affair occurring within the muscles and peripheral nerves, rendering any invocation of central, cerebral mental representations completely superfluous.
However, this classical associationist architecture suffered from severe systemic limitations when confronted with the realities of complex animal navigation. Watsonian peripheralism was thoroughly brittle; it could not adequately account for how animals adapted to sudden environmental obstacles, how they bypassed blocked alleys without engaging in random trial and error, or how they could navigate environments successfully when their sensory inputs were surgically or experimentally altered. By confining psychology to a strict peripheral input-output reflexology, classical behaviorism created an artificial, mechanized caricature of biological organisms that was increasingly out of step with naturalistic observation and emerging experimental evidence regarding behavioral plasticity.
2.2 Thorndike’s Law of Effect and the Drive-Reduction Model
Coinciding with Watson’s polemical assault was the deeply influential, mechanically oriented functionalism of Edward L. Thorndike. Through his pioneering nineteenth-century puzzle-box experiments with felines, Thorndike formulated the foundational operational axiom of early twentieth-century learning theory: the Law of Effect. Thorndike posited that when an arbitrary motor action produced a “satisfying state of affairs,” the neural connection between the situation (stimulus) and the action (response) was automatically and mechanically strengthened or “stamped in.” Conversely, actions resulting in discomfort or failure were stamped out. Central to Thorndike’s formulation was the absolute necessity of immediate consequences: learning was an automatic, blind process wherein reinforcement acted directly upon the neural connection, requiring no cognitive awareness, expectation, or comprehension on the part of the animal.
This mechanistic paradigm was subsequently formalized, systematized, and elevated to its axiomatic zenith by Clark L. Hull at Yale University. Throughout the 1930s and 1940s, Hull constructed an ambitious, mathematico-deductive learning theory centered upon the drive-reduction hypothesis. Hull postulated that learning occurs if and only if an organism’s action culminates in the reduction of an internal physiological drive state, such as hunger, thirst, or avoidance of pain. In the formal architecture of Hullian theory, the core unit of learning was habit strength, algebraically designated as $sHr$. The magnitude of habit strength was modeled as a direct, strictly monotonic function of the number of reinforced pairings between a stimulus and a response:
$$sHr = 1 – 10^{-aN}$$
Where $N$ represents the cumulative number of reinforced trials, and $a$ is an empirical constant. Within this hyper-formalized system, unreinforced experience was assumed to have an associative value of zero. If no primary drive reduction occurred, no reinforcement was delivered, and therefore habit strength could not increase. Hull’s model established an ironclad consensus throughout American experimental psychology: learning was theoretically impossible in the absence of explicit reinforcement. This conceptualization completely conflated the acquisition of knowledge with the presence of rewards, creating an ideological hegemony that viewed all biological learning through the reductive lens of physiological gratification.
2.3 Logical Positivism, Operationalism, and Tolman’s Neobehaviorism
During the 1930s, American experimental psychology sought to anchor its scientific credentials by adopting the epistemology of logical positivism, which had emerged from the Vienna Circle (led by thinkers such as Moritz Schlick and Rudolf Carnap), alongside the operationalism advanced by physicist Percy Williams Bridgman. Logical positivism maintained that a scientific proposition is meaningful only if it is either analytically tautological or empirically verifiable through direct sensory observation. Bridgman’s operationalism took this a step further, asserting that a scientific concept is synonymous with the corresponding set of operations used to measure it. For many behaviorists, this philosophical zeitgeist served as a justification to outlaw all unobservable theoretical constructs from psychological discourse.
Tolman, however, engaged with logical positivism and operationalism in an extraordinarily sophisticated and subversively progressive manner. Rather than using operationalism to strip psychology of its cognitive breadth, Tolman utilized it to legitimize unobservable cognitive constructs within a rigorous, objective behavioral science. Tolman recognized, as did the more advanced logical positivists, that theoretical science routinely makes use of unobservable entities—such as electrons, magnetic fields, and gravitational potentials—provided that these unobservables are rigorously, mathematically, and operationally anchored to measurable antecedents (independent variables) and observable consequences (dependent variables).
Tolman termed these operationally tethered internal states intervening variables. By defining intervening variables as functional entities that mediate between incoming environmental stimuli and outgoing behavioral responses, Tolman demonstrated that one could discuss internal representations, expectations, and cognitive maps without lapsing into unscientific mysticism or subjective introspectionism. An internal construct was fully valid if its operational parameters were systematically defined by experimental inputs (such as the number of hours of food deprivation or the geometric structure of a maze) and experimental outputs (such as traversal speed or error frequencies at choice points). This methodological brilliance allowed Tolman to build an unassailable neobehaviorist framework that satisfied the strictest criteria of logical empiricism while simultaneously rescuing cognitive processes from the anti-mentalist dogmatism of Watson and Hull.
3. Conceptualizing Latent Learning: Definitions, Distinctions, and Theoretical Premises
3.1 Defining Latent Learning
The primary empirical weapon that Tolman wielded to shatter the Thorndikian and Hullian consensus was the phenomenon of latent learning. Conceptually, latent learning refers to the acquisition of environmental knowledge, structural relations, and behavioral possibilities that occurs without obvious primary reinforcement and without an immediate, observable transformation in overt performance. Under latent learning conditions, an organism explores an environment in a state of apparent neutrality—experiencing no drive reduction, receiving no terminal reward, and exhibiting no obvious improvement in its behavioral efficiency. To the casual, classical behaviorist observer, the animal appears to be wandering aimlessly, learning nothing.
However, Tolman posited that during this period of unrewarded exposure, the organism is not idle. It is actively and continuously taking in information about the topological features, spatial pathways, dead ends, and relative locations of the environment. This acquired information is synthesized and stored in dormant or latent cognitive structures. The knowledge remains completely hidden from the experimenter’s view until a motivational incentive is introduced. Once a salient reward is placed within the environment, the animal rapidly accesses and deploys this latent knowledge base, immediately exhibiting high levels of behavioral proficiency that completely bypass the slow, incremental curves mandated by classical trial-and-error conditioning.
Latent learning posed a fundamental threat to the prevailing behaviorist dogma because it decoupled learning from external reinforcement. In classical S-R models, reinforcement was the literal motor of learning—the chemical or electrical glue that stamped associative bonds into the central nervous system. By demonstrating that animals could achieve profound spatial and relational mastery of complex physical mazes merely through passive, unrewarded exploration, Tolman proved that reinforcement was not an indispensable requirement for the acquisition of knowledge. The organism was revealed to be an autonomous, epistemologically proactive agent driven by an inherent tendency to map its world, rather than a passive automaton awaiting drive-reducing rewards.
3.2 The Epistemological Separation of Learning from Performance
The empirical demonstration of latent learning required a revolutionary epistemological distinction that continues to serve as a cornerstone of modern cognitive psychology: the categorical separation of learning from performance. Throughout the early history of behaviorism, these two concepts were treated as functionally synonymous. Because classical behaviorists rejected internal states, an animal’s observable performance was considered a direct, unmediated readout of its underlying associative habit strength. If an animal’s running speed did not increase, or if its errors at choice points did not decline over successive trials, behaviorists concluded with absolute certainty that no learning had occurred.
Tolman recognized this conflation as an egregious category mistake. He formally redefined learning as an internal, latent cognitive state transformation—a quiet restructuring of the organism’s mental models and expectancies regarding environmental contingencies. Conversely, he defined performance as the overt, measurable translation of that acquired knowledge into physical motor execution. Performance is not a direct reflection of learning alone; it is a complex, multiplicative function of learning, motivational states, current physiological drive levels, and situational incentives. An animal may possess a highly accurate, comprehensive internal map of an environment, yet it will show no outward manifestation of this knowledge if it lacks a compelling functional reason to traverse the space efficiently.
In this conceptual paradigm, motivation does not act as the mechanistic instrument that creates or stamps in learning. Instead, motivation operates as a selective activator or retrieval engine for performance. Drives such as hunger or thirst determine whether the latent cognitive structures will be translated into observable action. The systemic vulnerability of classical S-R models lay precisely in their inability to accommodate this duality. Because Hullian and Watsonian theories tethered associative formation directly to behavioral reinforcement and observable execution, the sudden, unreinforced mastery demonstrated in latent learning paradigms revealed an unresolvable blind spot in classical associationism.
3.3 Precursors and Early Observations of Latent Learning
While Tolman provided the comprehensive theoretical and empirical framework for latent learning, the initial empirical spark emerged from preliminary experiments conducted in his Berkeley laboratory by his doctoral student, Hugh Carlton Blodgett. In 1929, Blodgett published a pioneering study titled “The Effect of the Introduction of Reward upon the Maze Performance of Rats.” Blodgett observed that when rats were allowed to navigate a complex, multi-unit maze without receiving any food reward upon reaching the exit, their performance curves showed minimal improvement; they continued to make numerous errors and moved through the corridors with sluggish latency. However, as soon as food was introduced at the end of the maze after several days of non-rewarded exposure, these rats demonstrated an immediate, discontinuous drop in errors within a single trial, rapidly matching the proficiency of control rats that had been continuously reinforced from the start.
Blodgett’s findings presented a profound theoretical crisis to Hullian habit strength theory. If learning strictly required drive reduction, the unrewarded rats should have accumulated no habit strength ($sHr$) across their early trials. Their performance upon receiving food should have mirrored the slow, gradual learning curve of completely naive animals on their second day of training. Instead, their instantaneous mastery indicated that they had acquired substantial, fully formed knowledge during those non-rewarded trials. Blodgett characterized this phenomenon by noting that the rats had acquired the layout of the maze latently, yet lacked the motivational impetus to display it.
Tolman immediately recognized the monumental theoretical implications of Blodgett’s preliminary findings. Blodgett had uncovered empirical proof that an intrinsic, spontaneous tendency to explore physical environments exists in biological organisms—an exploratory drive that operates independently of homeostatic deficits like starvation. Tolman realized that this exploratory drive functioned as an epistemological mechanism: it motivated rodents to acquire information about their surrounding topology simply because the information existed to be mapped. Sensing an opportunity to deliver a definitive blow to the mechanistic S-R establishment, Tolman set out to construct a far more elaborate, methodologically immaculate, and quantitatively unassailable experimental program to prove the reality of latent learning beyond a shadow of a doubt.
4. The Seminal 1930 Tolman and Honzik Maze Experiments: Methodology and Evidence
4.1 Experimental Architecture and Design
In 1930, Edward C. Tolman and Charles H. Honzik published what would become one of the most famous and influential empirical papers in the history of experimental psychology: “Introduction and Removal of Reward, and Maze Performance in Rats.” Designed to eliminate every conceivable methodological vulnerability found in Blodgett’s earlier study, the Tolman and Honzik paradigm utilized a sophisticated 14-unit T-maze apparatus. This standardized apparatus was an architectural masterpiece of behavioral engineering: it consisted of a series of uniform, identical choice points where navigating subjects were repeatedly forced to choose between a true pathway toward the goal and an intricately designed, systematic blind alley. Mechanical doors were installed throughout the corridors to prevent rats from backtracking, ensuring that errors could be rigorously, objectively operationalized and quantified as full-body entries into blind alleys.
To rigorously evaluate the role of reinforcement in learning versus performance, Tolman and Honzik established three distinct experimental cohorts of rats, each subjected to precise nutritional maintenance schedules and identical handling protocols:
- Cohort 1: The Continuously Rewarded Control Group (CR). These animals received an immediate, highly palatable food reward upon reaching the goal box at the terminus of the maze on every single daily trial throughout the entire duration of the study. This group was designed to establish the baseline classical learning curve under conditions of continuous reinforcement.
- Cohort 2: The Continuously Non-Rewarded Control Group (CNR). These animals received absolutely no food reward upon reaching the goal box throughout the experimental period. Upon reaching the terminus, they were simply retained in an empty goal box for a brief interval before being removed. This group established the baseline performance curve of exploration in the perpetual absence of primary drive reduction.
- Cohort 3: The Experimental Delayed-Reward Group (DR). This was the critical experimental cohort. For the first ten consecutive days of the experiment, these rats were treated identically to the non-rewarded control group: they traversed the maze without receiving any food reward upon entering the goal box, wandering the corridors under unreinforced conditions. On Day 11, however, the experimental condition was abruptly altered: an immediate food reward was introduced into the goal box and maintained on all subsequent days.
4.2 Quantitative Findings and the Discontinuous Performance Curve
The quantitative results of the 1930 Tolman and Honzik experiment provided some of the most striking empirical curves in the history of behavioral research. From Day 1 to Day 10, the behavioral data conformed superficially to classical behaviorist predictions. The continuously rewarded control group (Cohort 1) displayed a standard, incremental, negatively accelerating learning curve: their average error rates dropped steadily from roughly ten errors per run to fewer than three by Day 10, accompanied by a corresponding decrease in traversal latency. In stark contrast, both the continuously non-rewarded control group (Cohort 2) and the delayed-reward experimental group (Cohort 3) demonstrated minimal overt progress. Their error curves declined only slightly—reflecting general maze habituation and the extinction of initial freezing behaviors—hovering between six and eight errors per run through Day 10.
The theoretical earthquake occurred on Day 12—the run immediately following the introduction of the food reward to Cohort 3. If reinforcement were the necessary antecedent for the acquisition of associative connections, Cohort 3 should have behaved on Day 12 like naive rats on their second day of training, displaying a modest, incremental drop in errors. Instead, Cohort 3 exhibited a discontinuous, precipitous collapse in error rates. In a single trial, their error count plummeted from an elevated baseline down to an average that was not merely equal to, but actually lower than the error rate of the continuously rewarded control group that had been reinforced for eleven consecutive days.
The trajectory of this discontinuous performance curve was undeniable. Over Days 12 through 17, the delayed-reward group maintained this exceptional level of navigational mastery, traversing the 14-unit T-maze with near-zero errors and lightning speed. Statistical analyses confirmed that the instantaneous mastery achieved by the delayed-reward cohort after just one reinforced exposure matched or surpassed the proficiency achieved by Cohort 1 through days of continuous reinforcement. The rats had not begun learning on Day 11; they had already mastered the complex topology of the 14-unit maze over the preceding ten days of non-rewarded exposure. The introduction of the food reward simply provided the motivational catalyst that unlocked this vast repository of latent knowledge, transforming dormant internal information into flawless motor execution.
4.3 Theoretical Implications of the Tolman-Honzik Paradigm
The empirical findings of the Tolman-Honzik experiment dealt a devastating blow to the core axioms of classical and Hullian behaviorism. First and foremost, the study provided undeniable empirical refutation of the strict necessity of reinforcement for the acquisition of knowledge. The delayed-reward rats had successfully mapped fourteen intricate choice points and blind alleys in the absolute absence of primary drive reduction. The mechanical stamping-in of connections via reward—the theoretical bedrock of Thorndike’s Law of Effect and Hull’s habit strength—was definitively shown to be non-essential for biological learning.
Second, the Tolman-Honzik paradigm demonstrated that non-rewarded exploratory trials were actively generating structured spatial representations rather than random behavioral noise. If unreinforced exploration merely produced random associations or progressive fatigue, the sudden introduction of food should have resulted in immense confusion and high error rates as the animal struggled to extinguish bad habits. The immediate, near-zero error performance demonstrated that the rodents were actively processing spatial relationships, identifying cul-de-sacs, and integrating the physical geometry of the maze into a coherent internal schema.
Third, the experiment definitively established reinforcement as an incentive factor regulating performance rather than an absolute prerequisite for learning. Reinforcement acts as a pragmatic switchboard: it communicates to the animal which environmental locations currently hold survival relevance, prompting the animal to utilize its existing cognitive structures to execute goal-oriented actions. The immediate academic reception was marked by intense controversy; Hull and his disciples at Yale immediately launched counter-replications and theoretical rationalizations to salvage their mathematical frameworks. However, the Tolman-Honzik paradigm stood as an enduring empirical milestone that permanently established the validity of latent learning, fundamentally altering the trajectory of American psychology.
5. The Architecture of Cognitive Maps: Spatial Representation and Mental Modeling
5.1 Tolman’s 1948 Formulation: ‘Cognitive Maps in Rats and Men’
In 1948, Edward Tolman published his crowning intellectual achievement in the pages of the Psychological Review: a sweeping, brilliant theoretical synthesis titled “Cognitive Maps in Rats and Men.” In this monumental paper, Tolman looked back over two decades of experimental evidence from his Berkeley laboratory and formally introduced the theoretical construct of the cognitive map into the scientific lexicon. Tolman set out to construct an entirely new neurocognitive metaphor for the mammalian brain, directly attacking the mechanistic analogies that had dominated early twentieth-century neurophysiology.
Tolman explicitly contrasted what he termed the “telephone switchboard” metaphor of the brain with his own radical alternative: the “central control room” metaphor. Classical S-R behaviorism, Tolman argued, operated on the assumption that the brain functions merely as an enormous, passive telephone exchange. Incoming sensory stimuli ping the peripheral receptors, traveling along afferent pathways to the central switchboard, where they are mechanically cross-connected to efferent motor lines, triggering outgoing muscular contractions. In this switchboard paradigm, the organism is structurally incapable of internal synthesis, deliberation, or autonomous modeling; it is purely a conduit through which environmental energy is redirected into physical work.
Tolman counter-proposed that the central nervous system functions as an active, highly sophisticated central control room. In this control room, incoming sensory inputs are not simply hooked up to output switches. Instead, the incoming impulses are worked over, elaborated, synthesized, and integrated into an internal, cognitive map-like representation of the environment. This cognitive map is a complex, structural mental model of the physical terrain, indicating the routes, paths, environmental barriers, and relative spatial relationships between diverse locations. Crucially, Tolman did not restrict this construct to animal navigation. In the latter sections of the 1948 paper, he boldly extrapolated the cognitive map concept to human psychology, arguing that human social, political, emotional, and ideological behaviors are governed by internal cognitive maps—and that many of humanity’s greatest sociopolitical pathologies stem from dangerous, narrow distortions in these mental landscapes.
5.2 Structural Topology: Strip Maps versus Comprehensive Field Maps
A foundational theoretical contribution of Tolman’s 1948 paper was his taxonomy regarding the structural topology of cognitive maps. Tolman recognized that not all internal representations possess the same degree of richness, flexibility, or geometric fidelity. He posited that cognitive maps exist along a qualitative structural continuum ranging from narrow strip maps to broad, comprehensive field maps:
- Narrow Strip Maps: These representations are structurally rigid, unidimensional, and highly localized. A strip map functions essentially as a mental trace of a single, highly specific route. It encodes an unyielding sequence of environmental cues linked to immediate responses (e.g., “go straight, turn right at the post, proceed ten paces, turn left”). A strip map possesses virtually no awareness of surrounding topographical space; if the designated path is physically blocked or altered, the organism possesses no internal alternative vectors and is reduced to helpless vacillation or random trial-and-error behaviors.
- Broad Comprehensive Field Maps: In stark contrast, a comprehensive field map is a multidimensional, allocentric mental representation of the entire environmental space. It encodes not merely paths, but the global spatial and geometric relationships among environmental objects, landmarks, open territories, and goals. An organism equipped with a comprehensive field map understands where it is in relation to its destination regardless of its current heading or starting point. If the primary route is suddenly blocked, the organism instantly computes an alternative detour or selects a novel, untraveled shortcut, navigating with spatial grace and tactical economy.
Tolman systematically analyzed the psychological and physiological conditions that cause organisms to regress from flexible field maps to maladaptive strip maps. He identified four primary factors: severe brain damage, excessive or traumatic emotional stress, overly intense physiological drive states (such as acute starvation), and rigid, stereotypic training regimens. Tolman noted that when an animal or a human being is subjected to extreme frustrative stress or excessive drive intensity, their cognitive perceptual field undergoes severe narrowing. They fixate obsessively upon immediate, direct paths toward a goal, losing the cognitive flexibility required to step back, perceive broader environmental vectors, and execute intelligent detours. This insight bridged animal spatial navigation with clinical and social psychology, offering a profound explanatory framework for human neurosis, prejudice, and political fanatacism.
5.3 Cognitive Maps as Latent Relational Networks
Modern cognitive science and network theory view Tolman’s formulation of the cognitive map as one of the earliest descriptions of an internal relational network. A cognitive map is not a literal, static photograph or cartographic parchment stored inside the cranium; rather, it is a dynamic, multi-layered relational network that encodes metrics, vectors, topological distances, and semantic contingencies. Within this latent relational architecture, environmental locations exist as informational nodes, while pathways, traversable vectors, and functional affordances exist as dynamic edges connecting those nodes.
One of the most remarkable properties of this cognitive architecture is its capacity for continuous, dynamic updating through unrewarded exploration. As an animal traverses an environment, its internal network constantly incorporates novel topological features, updates distance metrics, and registers environmental transformations (such as the sudden appearance of a barrier or the relocation of an egress point) without requiring primary reinforcement to stamp in these structural revisions. The cognitive map is an active, running simulation of the physical world that can be interrogated offline by the central nervous system to predict the outcomes of hypothetical movements before those movements are physically committed to muscle tissue.
Furthermore, Tolman highlighted the extraordinary capacity of cognitive maps to integrate multiple distinct, fragmented spatial experiences into a single, unified, coherent internal schema. A rodent may explore Alley A and Alley B on different days, under varying motivational conditions, without ever having traveled directly between the two endpoints. Yet, within the latent relational network of the cognitive map, the metric coordinates of these pathways are reconciled within an allocentric frame of reference. This enables the organism to perform instant vector calculations, selecting novel, untraveled shortcuts that bridge disparate regions of the environment—a cognitive capability that lies entirely beyond the theoretical boundaries of single-pathway S-R associationism.
6. Place Learning vs. Response Learning: Deconstructing Mechanistic S-R Associations
6.1 The Epistemological Conflict: Muscle Memory vs. Spatial Knowledge
As the theoretical war between Clark Hull’s mechanistic behaviorism and Edward Tolman’s purposive behaviorism intensified through the 1940s, experimental psychology arrived at a definitive, critical crossroads. The entire debate crystallized around a singular, highly contentious experimental question: What does an animal actually learn when it navigates an environment? Does it learn a mechanistic sequence of internal motor movements (response learning), or does it learn the physical location of objects in external space (place learning)?
The Hullian perspective, firmly anchored in classical peripheralism, claimed that learning is fundamentally the acquisition of response habits. When an animal successfully reaches food in a maze, it has stamped in a chained sequence of kinesthetic and motoric responses. The rat does not know “where” the food is in an allocentric sense; rather, it knows how to execute a specific neuromuscular pattern—such as “run straight, turn right, run straight, turn left.” The animal’s internal world was viewed as an egocentric succession of muscle twitches. If the rat was placed in a different physical orientation, its stamped-in motor habits were predicted to compel it to execute the identical muscular turns, regardless of whether those turns led toward or away from the physical food supply.
Tolman, conversely, championed place learning. He argued that animals do not learn muscular turns; they learn spatial coordinates. They acquire knowledge about where environmental objects reside in an objective, allocentric coordinate system. The rat’s internal representation encodes the fact that food is located at a specific point in external three-dimensional space—for example, “at the north window” or “under the overhead light.” Therefore, if the animal is introduced to the environment from an entirely novel angle or starting location, its behavior will not be dictated by chained motor habits. Instead, the animal will dynamically calibrate its motor output to guide itself toward that invariant spatial place, executing whatever muscular movements—left, right, swimming, or crawling—are necessary to reach the cognitive destination.
6.2 The Tolman, Ritchie, and Kalish (1946) Cross-Maze Experiments
To settle this profound theoretical conflict once and for all, Tolman, along with his colleagues B. F. Ritchie and D. Kalish, published a landmark empirical study in 1946: “Studies in Spatial Learning. II. Place Learning versus Response Learning.” To pit the two competing theories against each other in an experimental crucible, the researchers designed the celebrated cross-maze apparatus. The cross-maze consisted of four corridors meeting at a central choice intersection, forming a plus-sign configuration with two distinct starting points (designated as South and North) and two distinct goal boxes (designated as East and West).
The experimental subjects were divided into two rigorously controlled groups:
- The Response-Learning Group: For this group, reinforcement was tied strictly to the execution of a specific motoric response. Regardless of whether a rat started from the North or South entrance, it was required to make the exact same physical turn (for instance, always execute a right turn at the central junction) to reach the food. Consequently, if the rat started from the South, turning right led to the East box where food was placed; if it started from the North, turning right led to the West box, requiring the food to be relocated accordingly. The physical location of the reward changed constantly relative to the room, but the required motor response remained perfectly invariant.
- The Place-Learning Group: For this group, reinforcement was tied strictly to an invariant spatial location in the room. The food reward was permanently positioned in one specific goal box (for instance, the East box under a distinctive room window). If the rat was launched from the South starting point, it had to turn right to reach the food. However, if it was launched from the North starting point, it had to turn left to reach the exact same spatial location. The motor responses varied continuously, but the physical destination remained entirely constant.
The experimental results yielded a decisive victory for Tolmanian theory. The place-learning group acquired the navigation task with astonishing speed, mastering the maze with minimal errors in an average of only a few trials. The response-learning group, on the other hand, struggled profoundly. They made rampant errors, required vastly more trials to achieve basic proficiency, and in many instances failed to master the task entirely. The rats demonstrated an innate, overwhelming cognitive predisposition to learn where the place was rather than which muscle to flex. Spatial knowledge decisively trumped motor memory, dealing an empirical blow to Hull’s peripheral muscle-habit framework from which it never fully recovered.
6.3 The Sunburst Maze and Spatial Shortcuts
To deliver an even more dramatic demonstration of the allocentric vector capabilities of cognitive maps, Tolman, Ritchie, and Kalish devised what has become universally known as the sunburst maze paradigm. The apparatus began with an initial training phase: rats were trained to traverse a deliberately tortuous, roundabout pathway to secure food. The path required the rat to emerge from a starting apparatus, cross a central circular table, enter an indirect, winding alleyway that traced an elongated series of right-angle detours, and eventually arrive at a goal box positioned at a specific spatial coordinate relative to the broader room cues.
Once the rats had thoroughly learned this indirect pathway, the critical test trial commenced. The original winding alleyway was completely removed and replaced with the “sunburst” configuration: a wide array of radial avenues radiating outwards from the central circular table like the rays of a sunburst, spanning a comprehensive 360-degree perimeter. The directly straight-ahead path was blocked by a barrier. According to classical S-R conditioning, the rats should have suffered immediate, massive behavioral breakdown. Having only been conditioned to make specific right and left turns within the old alley, the animals should have chosen the paths most perceptually or kinesthetically similar to the initial segment of the trained path—typically selecting avenues immediately adjacent to the blocked entrance.
The behavior of the rats was breathtaking in its cognitive sophistication. Instead of choosing avenues adjacent to the original path, a vast statistical plurality of the rats systematically selected Path 6—the precise radial spoke that pointed along an unblocked, direct spatial vector straight toward the physical location of the goal box, despite the fact that they had never traveled down that spoke in their lives. The animals had not learned a mere chain of turns; through their training, they had acquired an allocentric, metric understanding of the physical space. They knew the global coordinates of the starting table and the global coordinates of the food box. When given the radial freedom of the sunburst apparatus, they calculated a novel, untraveled shortcut across empty space. This direct shortcut behavior provided irrefutable empirical proof of spatial vector navigation mediated by an internal cognitive map.
7. Purposive Behaviorism: Goals, Expectations, and Sign-Gestalt Theory
7.1 The Tripartite Sign-Gestalt Configuration
To provide a rigorous theoretical infrastructure for how cognitive maps are structured and operationalized at the level of individual environmental encounters, Tolman formulated sign-gestalt theory. Borrowing deeply from the perceptual holism of Koffka and the topological field dynamics of Kurt Lewin, Tolman rejected the associationist doctrine that organisms bind isolated sensations (S) to motor impulses (R). Instead, he proposed that learning consists of the acquisition of organized, three-part cognitive structures termed sign-gestalts. A sign-gestalt is a holistic perceptual-cognitive unit composed of three fundamentally inseparable components:
- The Sign: The initial environmental stimulus, landmark, or sensory configuration encountered by the organism (e.g., a specific choice point in a maze, a distinctive visual pattern, or an auditory tone).
- The Significate: The expected environmental consequence, terminal goal object, or subsequent physical state that lies beyond the immediate horizon (e.g., a bowl of water, an empty alley, an electrical shock, or a food pellet).
- The Sign-Gestalt Expectation (Means-End Relation): The cognitive, relational belief or hypothesis asserting that if the organism executes a specific behavioral commerce or action with the Sign, it will lead reliably to the Significate.
In this tripartite formulation, learning is not the stamping-in of a mechanical response to an isolated stimulus. Rather, learning is the cognitive acquisition of meaning: the animal learns that this sign signifies that outcome via this particular route. The organism does not acquire habits; it acquires expectations. Tolman’s integration of Gestalt perceptual principles meant that these sign-gestalts were processed holistically. The animal perceives the choice point not as an array of disconnected photons striking its retina, but as a meaningful navigational node pregnant with relational possibilities. The sign points beyond itself to the significate, binding the animal’s perception of the present directly to its anticipation of the future.
7.2 Expectancy Theory and Cognitive Hypotheses
A core corollary of sign-gestalt theory was Tolman’s revolutionary proposition that animals do not navigate environments through random, passive conditioning, but rather through the active formulation and empirical testing of cognitive hypotheses. When a naive rat is first introduced to a novel maze, its behavioral patterns at choice points are not purely chaotic. Tolman and his students demonstrated through meticulous choice-point analysis that animals frequently adopt systematic, organized strategies: a rat might systematically test a “spatial hypothesis” (e.g., always choosing the left alley for several trials), followed by a “visual hypothesis” (e.g., always choosing the lighted door), until environmental feedback confirms or refutes the expectation.
This formulation elevated the rodent from a passive recipient of environmental conditioning to an active, scientific problem-solver that formulates rudimentary hypotheses, tests them against environmental contingencies, and updates its internal sign-gestalts accordingly. Crucially, this theoretical model generated powerful, empirically testable predictions regarding what should occur when an animal’s cognitive expectancies are systematically violated. If an animal merely learned S-R reflex habits, altering the nature of a terminal reward—while keeping its basic reinforcing properties intact—should produce minimal disruption to ongoing motor execution.
The empirical reality was strikingly demonstrated in a legendary 1928 experiment conducted by Otto Tinklepaugh, one of Tolman’s intellectual allies. Tinklepaugh trained macaque monkeys on a visual spatial tracking task where a desirable food reward (a piece of banana) was placed under one of two identical cups in full view of the monkey. The animal was then briefly obscured behind a screen before being permitted to choose a cup. Under normal conditions, the monkey immediately retrieved the banana and consumed it. However, on the critical experimental trials, the experimenter surreptitiously swapped the hidden banana for a piece of crisp lettuce—a food that the monkey would normally accept and consume under neutral conditions. When the screen was raised and the monkey selected the correct cup, it did not simply eat the lettuce. Upon uncovering the lettuce, the monkey stopped dead in its tracks, stared at the cup, rushed around the room, peered underneath the table, shrieked in distress, and refused to eat the food, displaying extreme behavioral disorganization. The monkey had not learned a blind motor habit to reach for the cup; it held a specific, highly detailed cognitive expectancy of a banana. The profound disruption upon encountering the unexpected outcome served as definitive empirical proof that internal cognitive representations directly dictate animal behavior.
7.3 Means-End Readiness and Cathexes
Tolman further enriched purposive behaviorism by developing a comprehensive taxonomy of cognitive-motivational constructs to explain how goals and environments interact. Central to this taxonomy was the concept of means-end readiness. Tolman defined means-end readiness as an organism’s persistent, underlying cognitive predisposition to recognize and utilize certain classes of environmental objects or geometric pathways as instruments to achieve specific ends. It is an organism’s generalized cognitive preparedness to find water in depressions, to seek shelter in dark crevices, or to navigate along continuous surfaces. Means-end readiness forms the structural bedrock upon which specific, real-time sign-gestalts are constructed during spatial exploration.
Closely aligned with this was Tolman’s adoption and operational redefinition of the psychoanalytic term cathexis. In Tolmanian psychology, a cathexis refers to the learned or innate cognitive association established between a specific internal drive state and specific environmental goal objects. An animal is not merely driven by an undifferentiated, abstract “hunger”; through experience, it develops positive cathexes toward specific foods (such as sunflower seeds or lab chow) and negative cathexes toward unpalatable or toxic substances. These cathexes dictate which environmental significates will possess functional value for the organism when a drive is activated, directly shaping the motivational valence across its cognitive map.
Furthermore, Tolman introduced the concepts of equivalence beliefs and field expectancies. Equivalence beliefs refer to the cognitive processes by which an organism comes to treat a secondary subgoal—such as a specific visual signpost, a intermediate chamber, or a token—as possessing the reinforcing and motivating properties of the primary goal object itself, effectively anticipating modern computational ideas of sub-goal decomposition. Field expectancies, meanwhile, represent the larger spatial routing structures within the cognitive map that guide the animal along optimal routes toward positive cathexes while actively steering it away from negative cathexes. Together, these constructs formed an extraordinarily rich, multi-tiered architecture that seamlessly integrated cognitive mapping, motivational dynamics, and perceptual processing into a unified purposive system.
8. Tolman’s Intervening Variables: Operationalizing the Organismic ‘O’ in S-O-R Psychology
8.1 The Genesis of the S-O-R (Stimulus-Organism-Response) Framework
The dominant paradigm of early twentieth-century behaviorism was aggressively dyadic: it posited a direct, unmediated S-R (Stimulus-Response) reflex arc. In this radical input-output model, the physical organism was treated essentially as an epistemological black box—an empty, translucent conduit through which environmental stimuli elicited immediate muscular or glandular reactions. Watson had insisted that speculating about the internal contents of this black box was fundamentally unscientific, arguing that any invocation of intermediate organismic processes inevitably reintroduced Cartesian dualism and mentalistic superstition into psychology.
Tolman recognized that the dyadic S-R formula was mathematically and empirically incapable of explaining the profound behavioral plasticity, variability, and context-dependence displayed by biological organisms. Identical physical stimuli produce radically divergent behavioral responses depending upon the animal’s internal physiological state, its past environmental history, and its current navigational hypotheses. To resolve this inadequacy without sacrificing empirical rigor, Tolman pioneered the triadic S-O-R (Stimulus-Organism-Response) framework. He inserted the Organism—designated as the capital letter ‘O’—directly between the stimulus and the response, transforming the animal into an active, mediating processing center.
Tolman’s supreme methodological triumph lay in demonstrating that the internal states of the organism could be operationalized with complete scientific legitimacy. Drawing upon the principles of logical empiricism, Tolman insisted that these intermediate organismic processes—which he designated as intervening variables—must never be treated as untestable, subjective introspections. Instead, an intervening variable was defined as an abstract, objective functional construct that is mathematically and logically anchored to directly observable environmental inputs (the independent variables) on one side, and directly measurable behavioral outputs (the dependent variables) on the other. By rigorously tethering the internal ‘O’ to empirical operational definitions, Tolman legitimized the scientific study of mind while remaining firmly within the boundaries of objective behavioral science.
8.2 Taxonomy of Tolman’s Intervening Variables
To formalize this S-O-R architecture, Tolman developed a sophisticated taxonomy that categorized the factors governing behavior into three distinct, interrelated tiers: Independent Variables, Intervening Variables, and Dependent Variables. This taxonomic architecture formed a rigorous mathematical and conceptual model of purposive behavior:
- Independent Variables (Inputs): These represent the direct, measurable, and experimenter-controlled environmental and historical parameters. Tolman classified these into five primary categories:
- Environmental Stimulus Set ($S$): The physical geometry, lighting, odors, and topological configuration of the maze.
- Maintenance Schedule ($M$): The precise physiological deprivation state of the animal, operationalized by the number of hours of food or water deprivation, or the percentage of body weight reduction.
- Heredity ($H$): The genetic lineage, strain, and biological capacities of the animal.
- Age ($A$): The chronological and developmental stage of the organism.
- Prior Training ($T$): The cumulative history of prior trials, exposures, and environmental interactions.
- Intervening Variables (Internal Organismic Mediators): These are the theoretical, cognitive, and motivational constructs that operate within the organism. They cannot be observed directly with the naked eye, but they are logically and mathematically dictated by the configuration of the independent variables. They include:
- Cognitive Maps and Spatial Schemata: The internal structural models of the environment.
- Expectancies and Hypotheses: The anticipation of specific significates following specific signs.
- Demands and Drives: The operationalized motivational urges toward specific goal objects.
- Motor Skills and Capabilities: The physiological competence to execute physical locomotion.
- Dependent Variables (Outputs): These represent the objective, directly measurable behavioral actions executed by the organism in the physical world. These include running speed, latency to exit the start box, error frequencies at choice points, the number of turns into blind alleys, and traversal trajectory vectors.
In Tolman’s formal system, behavior ($B$) is expressed as an explicit mathematical function of these mediating intervening variables, which in turn are functions of the independent inputs:
$$B = f(I_1, I_2, I_3, …)$$
Where the intervening variables ($I$) are mathematically defined by:
$$I = g(S, M, H, A, T)$$
By establishing this rigorous functional dependence, Tolman demonstrated that cognitive maps, expectancies, and purpose were not fuzzy metaphysical concepts, but fully operationalized, quantitatively tractable components of a modern, predictive science of behavior.
8.3 Vicarious Trial and Error (VTE) as an Empirical Marker
A persistent challenge faced by Tolman was the demand from hostile S-R behaviorists for direct, observable behavioral evidence of internal cognitive processing occurring in real time. Mechanistic critics argued that intervening variables were merely theoretical fictions invented after the fact to explain an animal’s performance. Tolman’s brilliant empirical answer to this challenge was his documentation and operationalization of Vicarious Trial and Error (VTE).
When an animal approaches a critical choice point in a complex maze—an intersection where it must decide between two or more alternative avenues—it does not always execute an immediate, mechanical reflex turn. Instead, Tolman, along with his student K. F. Muenzinger, observed that the animal frequently pauses at the junction, arresting its forward locomotion. It then engages in a distinctive, alternating behavioral sequence: it swivels its head and forebody repeatedly back and forth, looking intently down Path A, then looking down Path B, swinging back toward Path A, and hesitating before finally committing to a single corridor. Muenzinger and Tolman coined the term Vicarious Trial and Error to describe these overt, vacillating head movements.
Tolman demonstrated that VTE was the direct physical manifestation of internal deliberation, cognitive hypothesis testing, and active mental calculation. The animal was not physically executing motor trials and errors down the corridors; rather, it was running those trials and errors vicariously inside its internal cognitive map. Crucially, Tolman provided compelling quantitative data regarding the temporal dynamics of VTE. Early in the learning process, VTE frequencies are low because the animal possesses no cognitive expectancies to evaluate. As learning proceeds and the animal begins constructing its cognitive map, VTE frequencies skyrocket, peaking precisely during the phase when hypotheses are being actively compared. Once the cognitive map is fully consolidated and navigation becomes automated, VTE frequencies decline sharply back to near-zero. VTE provided a stunning, observable behavioral window directly into the active, deliberative cognitive life of the navigating organism, delivering another powerful empirical validation of purposive behaviorism.
9. Methodological Innovations: The Evolution of Maze Paradigms and Experimental Control
9.1 Apparatus Engineering and Systematic Controls
The fierce empirical debates that engulfed experimental psychology during the 1930s and 1940s forced Edward Tolman and his research group to develop unprecedented standards of methodological rigor and apparatus engineering. Classical S-R critics, desperate to explain away latent learning and place learning without conceding cognitive representations, relentlessly postulated peripheral sensory artifacts. Critics argued that the rats were simply following accidental olfactory trails left by previous runs, reacting to minute tactile vibrations in the floorboards, or responding to subtle extra-maze visual cues that were imperceptible to human experimenters.
To systematically eliminate every conceivable peripheral sensory alternative, the Berkeley laboratory engineered remarkably sophisticated experimental controls. To neutralize olfactory cues, Tolman introduced interchangeable, rotatable maze segments. Alleys were constructed with modular wooden units that were systematically rotated, scrubbed, and chemically deep-cleaned between trials, ensuring that any residual scent trails would lead in contradictory directions or cancel out entirely. In many critical experiments, the maze floor was coated with heavy, uniform layers of sawdust that was vigorously mixed and raked after every single traversal to disperse any localized olfactory gradients.
To eliminate extra-maze visual cues, experiments were conducted within total light-elimination chambers, or inside standardized enclosures completely enclosed by heavy, light-absorbing black velvet curtains that extended from the ceiling to the floor, eliminating external room landmarks, windows, and light fixtures. When visual cues were deliberately tested, experimenters introduced artificial, systematically movable landmarks (such as distinctive illuminated cards or geometric shapes hung overhead) that could be mathematically rotated relative to the maze frame. Elevated mazes—narrow pathways suspended several feet in the air without side walls—were deployed alongside traditional enclosed mazes and water mazes, forcing animals to rely on pure allocentric spatial representation rather than physical wall-following strategies. This exhaustive methodological rigor systematically dismantled peripheral sensory explanations, forcing the scientific community to confront Tolman’s cognitive interpretations on their empirical merits.
9.2 The Use of Detour and Barrier Paradigms
Beyond traditional multi-unit T-mazes, Tolman and his associates pioneered the use of intricate detour and barrier paradigms—frequently referred to in the psychological literature as Umweg (detour) problems, a concept originally popularized by Wolfgang Köhler in his classic Gestalt studies of chimpanzee problem-solving. A detour problem is an experimental setup wherein the most direct physical path connecting the organism to its desired goal is physically obstructed by a transparent or impassable barrier, requiring the animal to turn its back on the goal and move away from it across an indirect, circuitous avenue in order to ultimately attain it.
In a classical S-R associationist framework, detour tasks present an almost insurmountable barrier to learning. Because the physical goal object (e.g., food) exerts an intense associative attraction, a naive animal is predicted to experience powerful, automatic reflex drives to push directly against the barrier. Moving physically away from the goal should have an associative valence of near-zero or negative, requiring long, painful sessions of trial-and-error extinction to overcome the animal’s natural forward vector. Tolman exposed the limitations of this mechanistic view by constructing complex detour mazes with multiple, interconnected bypass corridors of varying topological lengths.
Tolman’s experiments revealed that when a primary pathway was blocked at a specific junction, rats displayed immediate, insightful spatial planning. Rather than engaging in blind, frantic scratching at the obstruction, the animals paused, engaged in brief VTE behaviors, and immediately selected the most topologically economical detour that bypassed the obstruction. If a barrier was placed deep within a maze corridor, blocking both the primary path and an intermediate secondary path, the rats demonstrated a profound comprehension of the global layout: they did not waste energy venturing down the secondary path to its obstruction point, but immediately took an overarching tertiary detour that branched off far before the original junction. Tolman quantified planning latency, directional economy, and spatial detour efficiency, demonstrating that animals solve detour problems through an internal spatial simulation that assesses the topological accessibility of paths within their cognitive map.
9.3 Standardization of Motivational Deprivation States
Recognizing that motivational factors played a profound role in regulating performance and shaping the structural topology of cognitive maps, Tolman introduced rigorous standardization to the operationalization of motivational deprivation states. In early twentieth-century psychological experiments, motivational states were handled with shocking casualness; animals were simply described as “hungry” or “well-fed,” leading to massive, irreproducible behavioral variance across different laboratories. Tolman fundamentally transformed this practice by establishing strict, quantitative operational metrics for biological drives.
Tolman and his students operationalized hunger and thirst not by vague subjective impressions, but through precise percentage of baseline body weight reduction coupled with strictly timed maintenance schedules. Rats were weighed daily on precision balances and maintained at stable, designated percentages (typically 80% to 85%) of their free-feeding, ad-libitum body weight through carefully titrated post-trial feeding sessions. Furthermore, nutritional variables were rigorously controlled: animals were maintained on standardized, scientifically formulated diets to ensure that specific micronutrient deficiencies did not confound exploratory behaviors.
This quantitative precision enabled Tolman to make a profound, highly counterintuitive discovery regarding the interaction between motivation and cognition: the differential impact of high versus low drive states on the breadth of cognitive maps. While classical behaviorists assumed that higher drive states simply amplified habit strength and accelerated learning, Tolman demonstrated that extreme, acute drive states actually impaired the quality of cognitive representation. Under intense, desperate hunger or severe thirst, animals suffered from cognitive perceptual tunneling, constructing narrow, rigid strip maps that focused obsessively on immediate paths to reward. Conversely, under moderate, non-traumatic motivational conditions, animals explored their environments broadly and serenely, actively building comprehensive, highly detailed field maps with rich vector flexibility. Tolman’s methodological mastery revealed that true cognitive mapping requires a delicate balance between physiological drive and natural exploratory inclination.
10. Critical Debates and Epistemological Clashes: Clark Hull vs. Edward Tolman
10.1 The Mechanistic vs. Cognitive Neobehaviorist Schism
From roughly 1930 through the mid-1950s, American experimental psychology was thoroughly consumed by what historians of science now recognize as the great Tolman-Hull Debate. This was not merely an empirical dispute over maze protocols; it was an epic, high-stakes epistemological war over the ultimate nature of the psychological organism and the appropriate conceptual vocabulary of behavioral science. On one side stood Clark L. Hull and his formidable Yale school, representing a mechanistic, reductive, mathematico-deductive neobehaviorism. On the other stood Edward C. Tolman and his Berkeley school, championing a cognitive, holistic, purposive field theory.
Hull’s vision for psychology was explicitly Newtonian. He sought to construct a formal, axiomatic, highly mathematical deductive system modeled on the geometry of Euclid and the physics of Isaac Newton. Hull believed that all mammalian behavior could ultimately be predicted through a rigorous system of postulates, corollaries, and algebraic equations centered upon his famous formula for excitatory potential ($E$):
$$E = H \times D \times V \times K$$
Where excitatory potential ($E$) is the multiplicative product of habit strength ($H$), internal physiological drive ($D$), stimulus intensity dynamism ($V$), and incentive motivation ($K$). Hull’s system was unapologetically mechanistic; learning was viewed as the passive, deterministic accretion of S-R connections through drive reduction. There was no room in Hull’s mathematical universe for internal representations, mental maps, autonomous hypotheses, or qualitative meaning.
Tolman countered this mechanistic reductionism with his cognitive field theory, heavily influenced by the field physics of Albert Einstein and the psychological topology of Kurt Lewin. For Tolman, an organism navigating an environment was not a passive assemblage of habits governed by scalar algebraic equations, but an intelligent agent moving through a multidimensional psychological field. Learning was fundamentally the acquisition of a map, not a bundle of conditioned twitches. The debate forced a profound clash of scientific values: Hull offered extraordinary quantitative precision, formal axiomatic structure, and direct mathematical tractability; Tolman offered superior explanatory scope, psychological realism, and an open, flexible cognitive architecture that could gracefully accommodate complex behavioral plasticity.
10.2 The Hullian Counter-Explanations for Latent Learning
Confronted with the stunning empirical demonstrations of latent learning and place learning emerging from Berkeley, the Hullian camp launched a massive, decades-long counter-offensive designed to reframe Tolman’s cognitive findings within the orthodox mechanics of S-R drive reduction. The primary intellectual architect of this counter-reformation was Hull’s brilliant disciple, Kenneth W. Spence of the University of Iowa, who constructed an extraordinarily sophisticated neo-Hullian theoretical apparatus to explain away latent learning without conceding the existence of cognitive maps.
Spence advanced what became known as the fractional anticipatory goal response ($r_g – s_g$) hypothesis. Spence argued that when an animal traverses a maze, it engages in subtle, covert physiological responses—minute preparatory chewing, swallowing, or salivation movements—designated as fractional anticipatory goal responses ($r_g$). These internal bodily responses generate internal proprioceptive and kinesthetic stimuli, designated as $s_g$. Over repeated trials, these $r_g – s_g$ mechanisms become conditioned to environmental stimuli, forming a continuous, internal, purely physiological chain of secondary reinforcement that pulls the animal forward. Spence claimed that what Tolman called a “cognitive map” was nothing more than an intricate, covert chain of peripheral $r_g – s_g$ muscle twitches and proprioceptive feedback loops.
To explain latent learning, Hullians argued that the non-rewarded trials were not actually non-rewarded at all. They claimed that subtle, overlooked secondary reinforcers were operating constantly: the psychological relief of escaping the confined start box, the mild drive reduction achieved by reaching the larger goal box, the tactile sensations of running, and the curiosity-reducing effects of novel environmental stimuli. Furthermore, Hull and Spence engaged in an exhaustive “replication war,” publishing studies showing that if an animal’s sensory inputs were sufficiently degraded, or if the maze alleys were made completely homogeneous and devoid of distinctive cues, rats would fall back on pure response learning (muscle habits). For two decades, the psychological literature was flooded with hundreds of studies arguing over whether latent learning could or could not be demonstrated under varying, micro-manipulated experimental conditions.
10.3 Resolution and the Limits of Behaviorist Debate
By the mid-1950s, the great Tolman-Hull debate had gradually ground to an exhausting, theoretical impasse. Psychologists such as Edwin R. Guthrie, alongside an emerging generation of younger experimentalists, began to express deep disillusionment with both entrenched camps. The fundamental problem was that both paradigms had become so internally complex and topologically elastic that they could accommodate virtually any empirical outcome through post-hoc adjustments. If an experiment revealed sudden, unreinforced spatial mastery, Tolman hailed it as proof of a cognitive map, while Hullians invoked complex, unobservable networks of covert $r_g – s_g$ fractional anticipatory responses and secondary reinforcement. If an experiment revealed rigid, repetitive motor habits, Hullians proclaimed victory for S-R conditioning, while Tolmanians pointed to the narrowing of field maps into strip maps caused by excessive drive stress.
The realization dawned that the debate had ceased to be an empirical conflict and had devolved into a battle of unfalsifiable meta-theoretical languages. Both systems could explain the identical data point; they merely chose to describe it using radically different conceptual currencies. S-R theorists were forced to postulate increasingly arcane, invisible, internal proprioceptive mechanisms that were just as unobservable as the cognitive constructs they sought to eliminate, undermining their claim to pure peripheral operationalism.
This profound theoretical exhaustion played an instrumental role in paving the intellectual path toward the Cognitive Revolution of the late 1950s and 1960s. Younger psychologists—including George A. Miller, Jerome Bruner, and Noam Chomsky—grew thoroughly dissatisfied with the contortions required to maintain strict behaviorist vocabularies. They recognized that Tolman had been fundamentally correct all along: organisms possess internal, central cognitive architectures that represent, compute, and model the external world. Tolman’s purposive behaviorism served as the critical intellectual bridge across which experimental psychology marched, transitioning from the rigid mechanization of classical behaviorism to the rich, dynamic horizons of modern cognitive science.
11. Neuroscientific Validation: Hippocampal Place Cells, Grid Cells, and Spatial Navigation
11.1 The Discovery of Place Cells in the Hippocampus
For more than two decades following the publication of “Cognitive Maps in Rats and Men,” Tolman’s cognitive map remained a brilliant, theoretical intervening variable—an abstract construct validated by behavioral data, yet devoid of direct neurophysiological substrate. To the materialist critics of his era, the cognitive map was viewed as an epistemological phantom, a conceptual placeholder awaiting eventual reduction to simpler neural reflexes. That skepticism was permanently demolished in 1971 when neurophysiologists John O’Keefe and Jonathan Dostrovsky at University College London published a revolutionary empirical paper: “The Hippocampus as a Spatial Map. Preliminary Evidence from Unit Activity in the Freely-Moving Rat.”
Utilizing micro-electrodes implanted directly into the brains of freely navigating rodents, O’Keefe and Dostrovsky recorded the extracellular action potentials of individual pyramidal neurons within the CA1 region of the hippocampus. What they discovered was astonishing: specific, individual pyramidal neurons fired bursts of action potentials at high rates if and only if the animal was located within a specific, circumscribed physical territory within its testing enclosure. If the rat stepped outside this defined geographic zone, the neuron fell silent, while a completely different neighboring neuron burst into activity. O’Keefe coined the term place cells to describe these spatially tuned neurons, and the geographic area that activated a specific neuron was designated as that cell’s place field.
In 1978, John O’Keefe and Lynn Nadel published their monumental neuroscientific monograph, The Hippocampus as a Cognitive Map. Synthesizing hundreds of anatomical, physiological, and behavioral studies, O’Keefe and Nadel explicitly credited Edward C. Tolman as their direct intellectual forefather. They demonstrated that hippocampal place cell firing patterns are allocentric: a place cell fires when the animal occupies its spatial coordinate regardless of which direction the animal is facing, which motor movements it executes, or which route it took to arrive there. Place cell assemblies construct an internal, neural coordinate system that directly embodies Tolman’s allocentric cognitive map. Thirty-six years after Tolman coined the term, the physical, neuronal reality of the cognitive map was conclusively established, ultimately earning O’Keefe the 2014 Nobel Prize in Physiology or Medicine.
11.2 Grid Cells, Border Cells, and the Entorhinal Navigation System
The neurobiological validation of Tolman’s spatial architecture reached its apex in 2005 with the groundbreaking discoveries made by Edvard Moser and May-Britt Moser, working alongside their research team at the Norwegian University of Science and Technology. Recording from the medial entorhinal cortex (MEC)—a major cortical input structure feeding directly into the hippocampus—the Mosers discovered a profoundly sophisticated class of spatial neurons that they termed grid cells.
Unlike hippocampal place cells, which fire exclusively in a single localized environmental territory, an entorhinal grid cell fires at multiple regularly spaced locations as an animal traverses an open surface. What captivated the global scientific community was the breathtaking geometric organization of these firing fields: for any individual grid cell, its multiple firing fields tile the entire available two-dimensional physical environment in an exquisitely regular, periodic, hexagonal tessellation. The entorhinal cortex generates an intrinsic, highly organized internal metric coordinate system—a neural Euclidean grid—that measures physical distance, wavelength, orientation, and spatial scale across the terrain.
Subsequent neurophysiological investigations revealed that this entorhinal-hippocampal navigation circuit comprises an integrated, multi-layered global positioning system. Alongside grid cells and place cells, researchers identified:
- Head-Direction Cells: Located in the presubiculum and anterior thalamus, these neurons fire continuously whenever an animal’s head is oriented in a specific compass direction relative to the environment, functioning as an internal biological compass.
- Border Cells: Located in the medial entorhinal cortex, these cells fire exclusively when an animal approaches an impassable environmental boundary, wall, or geometric drop-off, providing physical boundary constraints to the cognitive map.
- Speed Cells: Neurons that modulate their firing rate linearly as a direct function of the organism’s physical running speed.
This interconnected circuitry executes real-time path integration (dead reckoning), continuously updating the organism’s internal spatial coordinates by integrating internal motion cues, velocity vectors, and direction signals. The entorhinal-hippocampal axis constitutes the exact neurobiological physicalization of the flexible, metric, comprehensive field map envisioned by Tolman in 1948, solidifying his status as a visionary prophet of modern cognitive neuroscience.
11.3 Modern Cognitive Neuroscience of Non-Spatial Mental Mapping
In contemporary cognitive neuroscience, the validation of Tolman’s vision has expanded far beyond the boundaries of physical terrain navigation. In the visionary closing passages of his 1948 paper, Tolman hypothesized that the cognitive mapping mechanism would ultimately be shown to govern human non-spatial thinking—including our mental modeling of social structures, semantic relationships, and abstract conceptual spaces. Over the past decade, cutting-edge functional magnetic resonance imaging (fMRI) and high-density intracranial recording studies in humans have definitively confirmed this extraordinary intuition.
Groundbreaking research conducted by cognitive neuroscientists such as Timothy Behrens, Christian Doeller, and Matthew Botvinick has revealed that the human hippocampal-entorhinal navigation system is actively co-opted to map abstract, multidimensional cognitive landscapes. When human subjects learn complex relational concepts—such as tracking intricate social hierarchies, navigating semantic networks, or categorizing novel multi-dimensional stimuli characterized by abstract parametric axes (e.g., varying neck lengths and leg lengths of artificial avian stimuli)—their brains deploy place-cell-like and grid-cell-like periodic activity within the entorhinal cortex and medial prefrontal cortex to navigate these non-spatial knowledge spaces.
The human brain constructs cognitive maps of conceptual space. Just as an entorhinal grid cell measures Euclidean physical distances in meters, it computes semantic and social distances between abstract concepts in high-dimensional feature space. This architecture enables humans to perform abstract inferential reasoning, compute conceptual shortcuts between seemingly disconnected ideas, and execute structural generalization—applying the structural geometry of a learned concept to an entirely novel intellectual domain. Tolman’s assertion that the brain is an active central control room that builds structural field maps has thus become the foundational paradigm across the whole of modern human neurobiology.
12. Enduring Legacy: Cognitive Psychology, Artificial Intelligence, and Modern Robotics
12.1 Catalyzing the Cognitive Revolution
Edward Chace Tolman’s intellectual insurgency served as the indispensable historical catalyst for the birth of modern cognitive psychology. Throughout the darkest decades of Watsonian and Skinnerian anti-mentalism, Tolman kept the flame of cognitive representation alive, demonstrating through unassailable empirical methodologies that an objective science of psychology could study internal models, hypotheses, and expectations without abandoning scientific rigor. He demonstrated to a skeptical world that cognitive science did not require unscientific introspection; it merely required brilliant experimental operationalism.
Tolman exerted an immense, direct personal and theoretical influence on the central architects of the Cognitive Revolution. Figures such as George A. Miller, whose classic work on information processing transformed psychology’s understanding of human capacity limitations, and Jerome Bruner, who pioneered cognitive constructivism and category formation, traced their intellectual lineages directly back to Tolman’s purposive neobehaviorism. Tolman’s intervening variables served as the direct conceptual prototype for the “information-processing stages,” “internal representations,” and “computational architectures” that came to define cognitive science in the latter half of the twentieth century.
The paradigm shift represented a definitive transition away from the passive associationism of classical mechanics toward active, predictive models of perception and cognition. Today, cognitive science views the brain as an active, prediction-generating organ—a predictive processing engine that continuously generates top-down generative models of the world, testing them against sensory prediction errors in a manner that directly mirrors Tolman’s sign-gestalt expectations and cognitive hypotheses. In every domain of contemporary cognitive inquiry, Tolman is celebrated not merely as a historical footnote who opposed Watson, but as the foundational conceptual forefather of modern cognitive science.
12.2 Reinforcement Learning in Artificial Intelligence
In modern computer science and artificial intelligence, Tolman’s theoretical concepts have undergone a breathtaking, high-tech renaissance. The mathematical framework of Reinforcement Learning (RL), which forms the computational backbone of modern autonomous agents, is explicitly divided into two major algorithmic paradigms that directly reflect the historical clash between Clark Hull and Edward Tolman:
- Model-Free Reinforcement Learning (The Thorndike/Hull Paradigm): In model-free RL (exemplified by classic algorithms such as Q-learning and Policy Gradient methods), an artificial agent learns by directly mapping sensory inputs (states) to motor outputs (actions) based purely on scalar reward values stamped into neural weights. The agent possesses no internal model of how the external environment actually operates; it does not know what the next state will be until it physically transitions there. It is a bundle of mathematical habits, computationally efficient but notoriously brittle, sample-inefficient, and incapable of adapting rapidly to environmental disruptions.
- Model-Based Reinforcement Learning (The Tolman Paradigm): In model-based RL, an agent constructs an explicit, internal mathematical world model (a digital cognitive map) that encodes state-transition dynamics: $T(s, a) \rightarrow s’$. The agent learns what the world looks like, how its actions transform the environment, and where rewards reside within that internal relational network. When presented with a task, a model-based agent does not blindly execute habit weights; it uses its world model to run internal, offline tree-search simulations—a direct computational implementation of Vicarious Trial and Error (VTE)—to plan multi-step trajectories, bypass newly introduced barriers, and execute novel shortcuts.
Furthermore, Tolmanian latent learning serves as the direct conceptual foundation for self-supervised and unsupervised representation learning in modern deep learning. Systems such as AlphaZero, MuZero, and modern world-model architectures (such as Hafner’s Dreamer networks) do not wait for explicit human rewards to begin understanding their domains. They engage in massive, self-supervised latent learning, exploring state spaces to construct rich, low-dimensional predictive representations of physical and game topologies before any task-specific reward is introduced. The deep learning revolution has demonstrated what Tolman proved in 1930: that autonomous biological or artificial intelligence requires the construction of an internal world model, and that latent learning is the ultimate engine of sample efficiency and structural generalization.
12.3 Simultaneous Localization and Mapping (SLAM) in Robotics
In the applied engineering realm of autonomous robotics, Tolman’s purposive behaviorism and cognitive mapping principles have become the practical engineering blueprint for machine navigation. Autonomous platforms—ranging from automated vacuum cleaners and industrial automated guided vehicles (AGVs) to planetary rovers and autonomous self-driving vehicles—cannot operate safely or effectively relying on simple S-R sensorimotor reflexes. A robotic car that merely executes a “turn right” reflex when a sensor detects an obstacle would be inherently catastrophic.
Modern autonomous robotics is fundamentally founded upon the algorithmic framework of Simultaneous Localization and Mapping (SLAM). SLAM algorithms address the identical computational dilemma that Tolman’s rats faced in the 14-unit T-maze: an autonomous agent placed in an unknown environment must incrementally construct a comprehensive, allocentric spatial map of the unfamiliar physical terrain (Mapping) while simultaneously tracking its own spatial coordinate location within that unfolding map (Localization). Utilizing probabilistic techniques such as Extended Kalman Filters (EKF), Particle Filters, and graph-based optimization (GraphSLAM), these autonomous agents construct detailed topological field maps that integrate sensor streams from LiDAR, visual cameras, and inertial measurement units.
Just as Tolman’s rats computed spatial vectors in the sunburst maze, modern SLAM systems perform continuous loop closure: recognizing when the robot has returned to a previously mapped spatial node, reconciling internal coordinate drifts, and updating the global geometric network. When an unexpected physical roadblock appears, a SLAM-guided autonomous vehicle does not freeze or rely on blind trial and error; it interrogates its internal topological map, runs instantaneous vector calculations, and executes an intelligent, real-time detour along an alternative pathway. Tolman’s purposive behaviorism has thus evolved from a controversial psychological thesis into the indispensable algorithmic foundation enabling machines to intelligently perceive, understand, and navigate the physical universe.
Conclusion: The Enduring Triumph of Purposive Cognition
Edward Chace Tolman’s lifelong intellectual odyssey constitutes one of the most transformative chapters in the annals of scientific psychology. Standing virtually alone against an anti-mentalist behaviorist orthodoxy that commanded American academia for half a century, Tolman refused to capitulate to the simplistic, mechanistic caricatures of mind and behavior that dominated his era. With extraordinary intellectual courage, profound philosophical nuance, and unassailable experimental ingenuity, he systematically dismantled the S-R reflex arc, demonstrating that biological life cannot be reduced to blind collections of stamped-in motor twitches.
Through the twin discoveries of latent learning and cognitive maps, Tolman fundamentally changed our understanding of the relationship between experience, knowledge, and action. He proved that learning is an autonomous, epistemologically proactive process driven by curiosity and an innate drive to map the world, existing independently of primary physiological reinforcement. By carving an indelible epistemological boundary between internal learning and overt performance, he liberated experimental psychology from the narrow confines of peripheralism, showing that the central nervous system is not a passive switchboard, but an active, predictive central control room that builds structural models of reality.
Today, as we witness the glorious convergence of cognitive psychology, systems neuroscience, and artificial intelligence, Tolman’s purposive behaviorism stands completely vindicated. From the hexagonal firing fields of entorhinal grid cells to the algorithmic world models powering cutting-edge autonomous robotics, the principles he uncovered in his humble Berkeley rodent mazes have revealed themselves to be fundamental, universal laws governing intelligent systems. Edward Chace Tolman did not merely rescue the mind from the dogmatism of classical behaviorism; he provided an enduring, magnificent intellectual blueprint for understanding the purposive, cognitive architecture of both biological and artificial minds.
References
- Blodgett, H. C. (1929). The effect of the introduction of reward upon the maze performance of rats. University of California Publications in Psychology, 4(8), 113–134. https://psycnet.apa.org/record/1930-01185-001
- Bridgman, P. W. (1927). The Logic of Modern Physics. Macmillan. https://archive.org/details/logicofmodernphy00brid
- Hafner, D., Lillicrap, T., Ba, J., & Norouzi, M. (2020). Dream to control: Learning behaviors by latent imagination. International Conference on Learning Representations (ICLR). https://arxiv.org/abs/1912.01603
- Hull, C. L. (1943). Principles of Behavior: An Introduction to Behavior Theory. Appleton-Century-Crofts. https://psycnet.apa.org/record/1943-03610-000
- Moser, E. I., Kropff, E., & Moser, M.-B. (2008). Place cells, grid cells, and the brain’s spatial representation system. Annual Review of Neuroscience, 31, 69–89. https://doi.org/10.1146/annurev.neuro.31.061307.090723
- O’Keefe, J., & Dostrovsky, J. (1971). The hippocampus as a spatial map: Preliminary evidence from unit activity in the freely-moving rat. Brain Research, 34(1), 171–175. https://doi.org/10.1016/0006-8993(71)90358-1
- O’Keefe, J., & Nadel, L. (1978). The Hippocampus as a Cognitive Map. Oxford University Press. https://global.oup.com/academic/product/the-hippocampus-as-a-cognitive-map-9780198572060
- Spence, K. W. (1936). The nature of discrimination learning in animals. Psychological Review, 43(5), 427–449. https://doi.org/10.1037/h0056975
- Thorndike, E. L. (1898). Animal intelligence: An experimental study of the associative processes in animals. The Psychological Review: Monograph Supplements, 2(4), i–109. https://doi.org/10.1037/h0092987
- Tinklepaugh, O. L. (1928). An experimental study of representative factors in monkeys. Journal of Comparative Psychology, 8(3), 197–236. https://doi.org/10.1037/h0075798
- Tolman, E. C. (1922). A new formula for behaviorism. Psychological Review, 29(1), 44–53. https://doi.org/10.1037/h0070289
- Tolman, E. C. (1932). Purposive Behavior in Animals and Men. Century Company. https://psycnet.apa.org/record/1932-04143-000
- Tolman, E. C. (1948). Cognitive maps in rats and men. Psychological Review, 55(4), 189–208. https://doi.org/10.1037/h0061626
- Tolman, E. C., & Honzik, C. H. (1930). Introduction and removal of reward, and maze performance in rats. University of California Publications in Psychology, 4(19), 257–275. https://psycnet.apa.org/record/1931-01646-001
- Tolman, E. C., Ritchie, B. F., & Kalish, D. (1946). Studies in spatial learning. II. Place learning versus response learning. Journal of Experimental Psychology, 36(3), 221–229. https://doi.org/10.1037/h0060262
- Watson, J. B. (1913). Psychology as the behaviorist views it. Psychological Review, 20(2), 158–177. https://doi.org/10.1037/h0074428