The study of instrumental conditioning in the mid-twentieth century was marked by a persistent tension between global reinforcement formulations and empirical anomalies that defied molar explanation. Classical paradigms derived from the systematic frameworks of Edward Thorndike and Clark L. Hull postulated that learned habits were forged through the direct, cumulative strengthening of stimulus-response (S-R) connections, mediated primarily by drive reduction. Within this framework, extinction—the cessation of responding when reinforcement is withdrawn—was envisioned as an orderly decay of associative strength or the progressive accumulation of reactive and conditioned inhibition. However, this molar conceptualization struggled to explain why organisms subjected to intermittent, unpredictable schedules of nonreinforcement frequently exhibited far greater resistance to extinction than those reinforced after every single behavioral execution—an enduring empirical challenge known as the Partial Reinforcement Extinction Effect (PREE).
Into this theoretical arena entered Egidio John Capaldi, whose sequential theory of instrumental learning revolutionized the behavioral sciences by redirecting analytical focus away from gross reinforcement percentages toward the precise, micro-structural sequence of individual trials. Capaldi proposed that nonreinforcement does not simply represent an absence of reward or a purely motivational insult; rather, it constitutes an active stimulus-generating event. Every trial outcome, whether reinforced or nonreinforced, leaves an internal, modified stimulus trace (or aftereffect) that persists across the intertrial interval to serve as a discriminative cue for subsequent behavior. By demonstrating that animals condition the instrumental approach response directly to the internal memory traces of nonreward, Capaldi constructed an objective, highly predictive associative architecture capable of explaining extinction persistence, sequential patterning, behavioral contrast, and nascent numerical competence without resorting to subjective mentalism or purely emotional constructs.
This treatise provides an exhaustive examination of Capaldi’s sequential theory, tracing its historical emergence, foundational tenets, mathematical formalizations, methodological innovations, and broader cognitive and neurobiological implications. Across decades of empirical work utilizing the discrete-trial straight-alley runway, Capaldi and his collaborators dismantled prevailing associative dogmas, illustrating that what an organism learns is intimately tied to the fine-grained serial architecture of its reinforcement history. By framing memory traces as functional intertrial stimuli, sequential theory bridged the divide between classical behaviorism and cognitive psychology, anticipating modern computational concepts of state representation, eligibility traces, and sequential decision-making in reinforcement learning.
1. Historical Context and Foundations of Sequential Learning
1.1 The Mid-Twentieth Century Neobehaviorist Landscape
The neobehaviorist landscape of the 1940s and 1950s was characterized by a concerted drive to formulate comprehensive, quantitative laws of behavior. Dominating this era was Clark L. Hull’s mathematico-deductive theory of learning, which formalized habit strength ($sHr$) as a monotonically increasing function of the number of reinforced pairings between an environmental stimulus ($S$) and an organismic response ($R$). In Hull’s formulation, drive reduction served as the primary mechanism through which associative bonds were cemented. Nonreinforced trials were treated not as sources of direct associative acquisition, but rather as inhibitory occurrences that fostered reactive inhibition ($I_R$) and conditioned inhibition ($sI_R$). Consequently, Hullian theory and its extensions by Kenneth Spence conceptualized learning through a molar lens: cumulative reinforcement strengthened approach habits, while extinction represented the unmasking of accumulated inhibitory potentials.
This mechanistic equilibrium was profoundly destabilized by empirical observations surrounding intermittent reinforcement schedules. If associative strength accrued solely through reward-driven drive reduction, then continuous reinforcement (CRF)—in which every correct response yields a reinforcer—should logically generate the maximum possible habit strength and, consequently, the greatest persistence when reward is discontinued. Yet, experimental laboratories consistently revealed the inverse: organisms trained under partial reinforcement schedules (PRF) persisted in running down alleys or depressing levers far longer during extinction than their continuously reinforced counterparts. This anomaly, which came to be known as the Partial Reinforcement Extinction Effect, struck at the core of drive-reduction mechanics, suggesting an alarming disconnect between total reinforcement quantity and behavioral persistence.
In response to these anomalies, behavioral science began splintering into competing camps. Cognitive theorists, inspired by Edward C. Tolman, posited that animals construct cognitive maps, field expectancies, and subjective hypotheses regarding environmental contingencies. From the Tolmanian viewpoint, extinction in CRF animals was rapid because the immediate shift to nonreinforcement generated an easily detectable discrepancy between expectation and reality; conversely, PRF animals persisted because the transition to nonreward was ambiguous. While intuitively appealing, cognitive formulations were frequently critiqued by rigorous neobehaviorists for their lack of operational precision, mathematical tractability, and vulnerability to homuncular explanations. It was precisely this conceptual deadlock that E.J. Capaldi sought to transcend, recognizing that what was needed was not a retreat into non-mechanistic mentalism, but an objective, molecular, stimulus-based account that incorporated the trial-by-trial sequential order of events into a rigorous associative framework.
1.2 Precursors to Sequential Formulations
Capaldi’s intellectual predecessors had glimpsed the importance of trial-level transitions, yet they lacked the systematic methodology necessary to formalize them into a predictive model. Clark Hull had recognized that sensory inputs do not vanish instantaneously upon physical offset; rather, they initiate a perseverative stimulus trace ($s$) that rapidly rises to a peak and then gradually dissipates across time. Hull utilized this concept primarily to bridge the temporal gap between the onset of a conditioned stimulus and the delivery of the unconditioned stimulus in classical conditioning, or between response execution and delayed reinforcement. However, Hull failed to apply this mechanism to the post-trial consequences of nonreinforcement, assuming that the cessation of reward left behind only an absence of stimulation rather than a distinctive, durable cue.
A more proximate catalyst for sequential thinking was Fred D. Sheffield’s 1949 investigation into the conditioned aftereffects of nonreward. Sheffield suggested that nonreinforced trials induce physiological or motoric aftereffects that remain present when the animal is placed back into the apparatus for the next trial. If a rewarded trial immediately follows a nonrewarded trial, this residual aftereffect can theoretically become conditioned to the instrumental response. While pioneering, Sheffield’s hypothesis was largely qualitative, tied closely to peripheral motor postures, and failed to systematically track how differing lengths of nonreinforced runs or varying temporal intervals altered these internal cues. Sheffield ultimately did not construct a mathematical or comprehensive theoretical architecture capable of predicting extinction across diverse schedules.
The fundamental flaw of the early neobehaviorist paradigms lay in their reliance on molar reinforcement models, which collapsed experimental protocols into crude summary statistics, such as percentage of reinforcement (e.g., 50% vs. 100%) or total reward volume. By focusing exclusively on cumulative response counts, researchers ignored the micro-structural architecture of learning. Two experimental schedules delivering identical 50% reinforcement could possess vastly different serial organizations: one might alternate strictly between reward and nonreward, while another clustered nonreinforcements into extended runs. Capaldi perceived that organisms do not react to statistical summaries computed post-hoc by an experimenter; they encounter discrete, localized transitions in real time. The paradigm shift initiated by sequential theory required abandoning molar aggregates in favor of an exacting micro-analysis of reinforcement history.
1.3 E.J. Capaldi’s Methodological Departure
To capture the fine-grained dynamics of sequential learning, Capaldi instituted a methodological departure from the then-prevalent free-operant methodologies popularized by B.F. Skinner. In a free-operant chamber, an animal emits responses at its own rate, rendering the intertrial interval fluid, uncontrolled, and intrinsically confounded with the organism’s own vigor. Capaldi turned instead to the highly controlled, discrete-trial straight-alley runway paradigm. In this apparatus, the experimenter strictly governs the initiation of every trial, the precise temporal confinement within the goal box, and the duration of the intertrial interval (ITI). This experimental control allowed Capaldi to manipulate sequence architecture, reward magnitudes, and intertrial durations with millisecond precision, ensuring that internal aftereffects could be systematically correlated with subsequent behavioral performance.
Through this meticulous paradigm, Capaldi formulated a systematic stimulus-trace hypothesis grounded not in speculative cognitive inferences, but in observable physical manipulations. He asserted that nonreinforcement functions as an active stimulus event that generates a distinct, conditionable aftereffect. Rather than viewing the animal as a passive accumulator of habit strength, Capaldi conceptualized the subject as navigating a sequence of internally generated stimulus cues. Each trial outcome dynamically transformed the internal sensory environment encountered by the animal at the start of the subsequent trial. By varying the sequence of rewarded (R) and nonreinforced (N) trials, Capaldi could directly measure how specific sequence geometries governed the rate of runway acquisition, asymptotic running velocity, and resistance to extinction.
Crucially, Capaldi rejected purely emotional or motivational accounts of sequential phenomena, positioning his theory in direct opposition to contemporary frustration models. While others argued that nonreinforcement acted primarily as an affective shock producing visceral emotional reactions, Capaldi maintained an intensely cognitive yet behaviorally rigorous stance: nonreinforcement produces an informational, memorial cue. This internal stimulus trace obeys standard laws of associative conditioning, generalization, and temporal decay. By treating memory traces as internal discriminative stimuli, Capaldi stripped the analysis of intermittent reinforcement of subjective emotionalism, laying the groundwork for an elegant, stimulus-based architecture of instrumental learning.
2. Core Tenets of Capaldi’s Sequential Theory
2.1 The Construct of Modified Aftereffects (S^N and S^R)
At the center of Capaldi’s sequential formulation lies the construct of modified stimulus aftereffects. When an organism completes an instrumental response and arrives at the goal area, it encounters a specific outcome that alters its internal neurophysiological state. Capaldi designated the internal stimulus trace generated by the consumption of a reward as $S^R$, whereas the trace produced by the absence of reward (nonreinforcement) was designated as $S^N$. These are not mere passive sensory flickers; they are persistent internal stimuli possessing cue properties that survive the termination of the physical outcome and persist across the intertrial interval into the subsequent trial.
The internal aftereffect $S^N$ is conceived as a dynamic physiological and cognitive trace that undergoes systematic temporal decay. Immediately following the realization of nonreward in the goal box, $S^N$ exhibits its highest intensity. As the intertrial interval elapses, the intensity of $S^N$ decays according to a monotonic temporal gradient. Despite this attenuation, Capaldi demonstrated that with short to moderate ITIs, a significant portion of the $S^N$ trace persists into the subsequent trial. Consequently, when the animal is reintroduced into the runway start box on trial $t+1$, its internal sensory environment is comprised not only of external apparatus cues (the tactile, visual, and olfactory properties of the alleyway), but also of the lingering $S^N$ or $S^R$ trace originating from trial $t$.
Because $S^N$ and $S^R$ represent fundamentally disparate internal events, they possess distinct, differentiable cue properties. An organism can learn to use these internal aftereffects to guide behavior with exceptional precision. Just as an animal can learn to respond to an external auditory tone or visual light, it can learn to initiate an instrumental response in the presence of $S^N$ or $S^R$. The internal outcome trace ceases to be merely a retrospective record of past events; it is transformed into an active, prospective cue that enters into functional associative relationships with the instrumental response.
2.2 Associative Mechanisms Governing Outcome Transitions
The foundational associative mechanism of Capaldi’s theory operates upon outcome transitions across sequential trials. Instrumental conditioning typically focuses on the pairing of external cues with an instrumental motor response ($R_{inst}$) followed by a reinforcer. Sequential theory expands this formulation: the stimulus compound eliciting the instrumental response at the start of trial $t+1$ is represented as a complex configuration containing both the external apparatus cues ($S_{app}$) and the prevailing internal aftereffect ($S^N$ or $S^R$). Thus, the nominal stimulus complex is formalized as $(S_{app} + S^N)$ or $(S_{app} + S^R)$.
Consider an experimental sequence where a nonreinforced trial (N) is followed immediately by a reinforced trial (R)—a configuration denoted as an N-R transition. On trial 1 (N), the animal runs the alley, finds no food, and experiences nonreinforcement, generating the aftereffect $S^N$. On trial 2 (R), the animal is placed back into the alley while $S^N$ is still active. The animal traverses the runway and receives food in the goal box. In this critical moment, the instrumental approach response ($R_{app}$) executed in the presence of $S^N$ is rewarded. Through standard associative processes, the internal trace $S^N$ becomes directly conditioned to the instrumental response ($S^N \rightarrow R_{app}$). This single mechanism forms the bedrock of conditioned persistence.
When successive nonreinforced trials occur (e.g., an N-N-R sequence), the aftereffect of nonreinforcement undergoes compounded modification. The second nonreinforced trial occurs in the presence of the decaying trace from the first, generating a modified, higher-order internal trace. Capaldi formalized these compounding states, asserting that an animal subjected to multiple consecutive nonreinforcements does not experience a static internal environment, but rather an evolving series of modified traces. As these varied aftereffect states are repeatedly followed by reward, a broad generalization gradient is forged, anchoring the instrumental approach response across an entire multidimensional continuum of internal nonreward states.
2.3 The Dual Role of Trace Intensity and Trace Quality
Capaldi differentiated between the quantitative intensity and the qualitative identity of sequential aftereffects. The intensity of an aftereffect is primarily a function of nonreinforced sequence length and temporal decay. As an animal undergoes a sequence of consecutive nonreinforced trials, the cumulative impact of nonreinforcement alters the internal state, driving it along a dynamic continuum. A run of two consecutive nonreinforced trials produces a trace designated as $S^{N2}$, while three consecutive nonreinforcements yield $S^{N3}$. These numerical superscripts reflect both a quantitative shift in trace magnitude and an altered perceptual quality.
The qualitative dimension of an aftereffect allows organisms to discriminate between distinct magnitudes of nonreward aftereffects. A nonreward event following an expected large reward generates a quantitatively more intense and qualitatively different trace than a nonreward event following an expected small reward. Furthermore, trace interaction governs how these states interact: when an animal encounters a rewarded trial after varying runs of nonreward, the prevailing trace—whether $S^{N1}$, $S^{N2}$, or $S^{N3}$—enters into immediate competition and compounding with the reinforcing event. The animal’s internal representation becomes a complex vector space where previous outcomes modulate current associative processing.
To capture these interactions mathematically, Capaldi formulated retention and decay functions characterizing the transition of traces across discrete intervals. If $I(S^N_0)$ represents the initial intensity of the nonreward aftereffect at the moment of goal-box exit, the effective intensity at the onset of the subsequent trial, separated by an intertrial interval $t$, can be modeled as a decay function:
$$S^N(t) = S^N_0 \cdot e^{-\lambda t}$$
where $lambda$ represents an empirical decay parameter influenced by organismic variables, apparatus characteristics, and the depth of retroactive interference. By formalizing these dynamics, sequential theory demonstrated that internal traces behave with the same mathematical discipline as physical conditioned stimuli in traditional Pavlovian paradigms.
3. The Partial Reinforcement Extinction Effect (PREE)
3.1 The Paradox of Intermittent Reinforcement
The Partial Reinforcement Extinction Effect (PREE) represents one of the most robust and extensively investigated phenomena in comparative psychology. Early investigations by Humphreys (1939) and Skinner (1938) established that intermittent reward leads to extraordinary persistence when reinforcement is utterly eliminated. Under classical associative models, this observation constituted an insurmountable paradox. Hullian mechanics dictated that habit strength was a cumulative function of rewarded trials; therefore, an animal receiving 100% continuous reinforcement (CRF) across 100 trials must possess an associative habit strength far superior to an animal receiving only 50% partial reinforcement (PRF) across the same 100 trials. Under extinction, the CRF animal should theoretically possess more associative capital to deplete and, hence, should persist substantially longer. Reality flatly contradicted the model.
Traditional non-sequential attempts to resolve the paradox typically relied on postulating vague discriminative factors or hypothetical emotional mechanics. Capaldi demolished this paradox through an elegant associative resolution: extinction performance is directly determined by whether the stimuli present during extinction have been conditioned to the instrumental response during the acquisition phase. In extinction, every single trial is, by definition, a nonreinforced trial. Consequently, every extinction trial (from the second trial onward) occurs in the presence of $S^N$. For a continuously reinforced animal, $S^N$ was never once encountered during training; acquisition occurred exclusively in the presence of $S^R$ and neutral apparatus cues.
The inverse relationship between acquisition rate and extinction durability thus receives a clean, non-paradoxical explanation. CRF animals run rapidly during acquisition because their trials are free from the disruptive, unconditioned competing responses initially provoked by nonreward. However, when placed into extinction, they immediately encounter a novel stimulus state ($S^N$) to which approach responding has never been conditioned. Conversely, PRF animals acquire the instrumental habit more slowly because they must learn to navigate fluctuating internal cues, but they exhibit monumental persistence in extinction because the cues defining extinction are the very cues that previously signaled that reward was forthcoming.
3.2 The Role of N-R Transitions in PREE
A primary empirical breakthrough of Capaldi’s sequential theory was the identification and isolation of the N-R transition as the indispensable architectural unit driving the PREE. An N-R transition is formally defined as an experimental sequence wherein a nonreinforced trial (N) is followed immediately by a reinforced trial (R). Capaldi asserted that the partial reinforcement extinction effect does not emerge simply from the passive experience of nonreward, nor does it correlate strictly with the percentage of reinforced trials. Rather, PREE is driven by the explicit pairing of $S^N$ with an instrumental response that is subsequently reinforced.
To validate this foundational principle, Capaldi designed ingenious sequence-manipulation experiments that decoupled total reinforcement percentage from the presence of N-R transitions. In a classic experimental arrangement, two groups of rats were administered identical schedules containing exactly 50% reinforcement across an identical number of total trials. Group 1 received a schedule rich in N-R transitions (e.g., N-R-N-R-N-R). Group 2 received an identical number of N and R trials, but organized into blocked configurations such that all nonreinforced trials occurred together or transitions to reward were systematically eliminated (e.g., R-R-R-N-N-N). Despite possessing identical reinforcement percentages and identical cumulative experiences with nonreward, Group 1 displayed robust resistance to extinction, whereas Group 2 extinguished rapidly, demonstrating behavioral vulnerability equivalent to continuously reinforced animals.
These findings revealed that the cumulative frequency of N-R pairings during acquisition correlates directly with extinction resistance. Under Capaldi’s framework, each N-R transition serves as an explicit conditioning trial for the internal stimulus trace: $S^N$ enters into an associative bond with the approach response ($S^N \rightarrow R_{app}$). Through repeated N-R experiences, $S^N$ acquires discriminative control over the instrumental response. When the animal is subsequently plunged into extinction, the omnipresent $S^N$ traces do not provoke cessation; instead, they function as powerful conditioned elicitors of the approach response, driving the animal repeatedly down the runway despite the total absence of food.
3.3 Generalization Decrement as a Competing Hypothesis
To fully explain extinction behavior, Capaldi integrated the associative conditioning of $S^N$ with the formal principle of stimulus generalization decrement. Generalization decrement refers to the attenuation of conditioned responding that occurs when the physical or internal stimulus complex present during testing differs from the stimulus complex present during initial acquisition training. If an animal is conditioned to respond to a pure auditory tone of 1000 Hz, shifting the test tone to 800 Hz produces an immediate drop in response magnitude—a classic generalization deficit.
Capaldi applied this rigorous psychophysical logic directly to internal outcome aftereffects. For an organism subjected to a continuous reinforcement (CRF) schedule, every single acquisition trial begins in the presence of external apparatus cues plus the rewarding aftereffect: $(S_{app} + S^R)$. The animal has learned to approach exclusively within this specific stimulus framework. When the experimenter shifts the animal to extinction, Trial 1 is nonreinforced. By the time Trial 2 begins, the goal-box outcome of Trial 1 has populated the animal’s internal environment with $S^N$. The stimulus compound is abruptly transformed into $(S_{app} + S^N)$. Because $S^N$ is markedly dissimilar from $S^R$, the CRF animal experiences a catastrophic generalization decrement. The runway cues no longer match the internal conditioning context, leading to an immediate collapse of instrumental running behavior.
In contrast, the partially reinforced animal has spent its entire acquisition phase executing approach responses under both $(S_{app} + S^R)$ and $(S_{app} + S^N)$ conditions. When extinction arrives, the emergence of $S^N$ on Trial 2 does not generate an alien internal landscape; rather, it introduces a familiar, highly conditioned stimulus environment. The PRF animal suffers negligible generalization decrement because the stimulus compound characterizing the extinction phase is structurally and perceptually identical to the stimulus compounds that commanded reward throughout acquisition. Capaldi thus differentiated between simple associative transfer and stimulus generalization failure, showing that much of what had traditionally been attributed to “habit loss” during extinction was actually an acute instance of generalization decrement provoked by sudden internal stimulus shifts.
4. N-Length, Sequence Architecture, and Run Dynamics
4.1 The Significance of Maximum N-Length
Moving beyond simple N-R transitions, Capaldi’s sequential theory provided a granular account of sequential runs, establishing the construct of “N-length.” N-length is defined as the number of consecutive nonreinforced trials preceding a reinforced trial. For instance, in the sequence R-N-R, the N-length is 1 (designated as $N_1$). In the sequence R-N-N-R, the run contains two consecutive nonreinforcements, yielding an N-length of 2 ($N_2$). If an animal experiences five nonreinforcements prior to reward (R-N-N-N-N-N-R), the sequence constitutes an $N_5$ run. Capaldi established that the maximum N-length experienced during acquisition acts as a decisive determinant of subsequent extinction persistence.
The theoretical rationale for the power of maximum N-length is rooted in trace differentiation. As nonreinforced trials accumulate in an unbroken string, the internal sensory environment undergoes systematic progression: $S^{N1}$ gives way to $S^{N2}$, which transforms into $S^{N3}$, and so on. If an animal is trained exclusively on an $N_1$ schedule (alternating R-N-R-N), the approach response is conditioned solely to the aftereffect of a single nonreinforced trial ($S^{N1}$). When this animal enters extinction, it performs robustly across the first few trials. However, as nonreinforced trials continue unabated through Trials 3, 4, and 5, the internal trace inevitably compounds into $S^{N3}$, $S^{N4}$, and $S^{N5}$. Because the $N_1$-trained animal has never experienced reward in the presence of these higher-order compound traces, it experiences progressive generalization decrement and rapidly halts.
Conversely, when an acquisition schedule incorporates long N-lengths—such as an $N_4$ or $N_5$ run—the animal receives reinforcement at the termination of the run, successfully conditioning the approach response to these severe, compound nonreward traces ($S^{N5} \rightarrow R_{app}$). By training the animal to persist through deep, extended runs of nonreward, the experimenter extends the associative umbrella across the entire spectrum of aftereffects that will inevitably arise during prolonged extinction. Capaldi proved empirically that matching the maximum N-length between acquisition and extinction virtually abolishes generalization decrement, establishing unprecedented levels of instrumental persistence.
4.2 Patterned versus Random Partial Reinforcement Schedules
The analytical power of sequential theory is nowhere more vividly demonstrated than in the investigation of patterned reinforcement schedules. A classic patterned paradigm is the single-alternation schedule, wherein reinforced and nonreinforced trials alternate with absolute predictability: R-N-R-N-R-N. Under this regimen, every R trial is preceded by an N trial (an N-R transition), and every N trial is preceded by an R trial (an R-N transition). If animals were merely insensitive statistical aggregators of a 50% reinforcement baseline, their running speeds across trials would remain uniform. Instead, Capaldi and his contemporaries revealed that animals develop exquisite “patterned running.”
After sufficient exposure to a single-alternation schedule, rats display rapid running speeds on R trials and pronounced deceleration or hesitation on N trials. Sequential theory explains this phenomenon: on every R trial, the prevailing stimulus cue in the start box is $S^N$ (derived from the preceding N trial); on every N trial, the prevailing cue is $S^R$ (derived from the preceding R trial). Because reward is consistently delivered in the presence of $S^N$ and consistently withheld in the presence of $S^R$, the animal forms a precise discriminative habit. The internal trace $S^N$ becomes an excitatory discriminative stimulus ($S^+$) signaling reward, while $S^R$ becomes an inhibitory discriminative stimulus ($S^-$) signaling nonreinforcement. The animal runs fast when it “remembers” nonreward and runs slow when it “remembers” reward.
This patterned performance represents an empirical refutation of the idea that PRF effects are governed purely by random stochastic persistence. It demonstrates that animals form discriminative sequential expectations based on internal outcome memories. If the schedule is suddenly shifted from single-alternation to a random sequence, the patterning disintegrates as the strict correlation between specific aftereffects and subsequent outcomes is shattered. The transition from patterned discrimination to generalized persistence under random schedules directly reflects the broadening of associative conditioning across diverse aftereffect states, proving that the animal’s internal memory states act as functional cues possessing direct behavioral control.
4.3 Distribution and Spacing of N-R Transitions
Sequence architecture involves not merely the presence of N-R transitions, but their precise spatial and temporal distribution across the learning curve. Capaldi demonstrated that massing versus distributing N-R transitions exerts profound effects on acquisition dynamics and memory stabilization. When N-R transitions are massed early in acquisition, the conditioning of $S^N$ occurs before the baseline habit strength of the instrumental response has fully stabilized, often leading to slower overall acquisition rates but resilient long-term behavioral persistence. Conversely, distributing N-R transitions across the entire training phase allows approach habits to intertwine continuously with shifting internal traces, fostering maximal transfer.
Furthermore, Capaldi investigated the structural impact of early versus late introduction of N-length variations. If an organism is exposed to long N-lengths early in training, the intense nonreward traces are introduced when associative habit strength is fragile, which can elicit profound behavioral suppression or avoidance. However, if the experimental sequence introduces escalating N-lengths progressively—graduating from $N_1$ to $N_3$ and ultimately $N_5$—the animal smoothly transfers approach responding across increasingly distant aftereffects. This graded sequence exposure builds structural resilience, systematically insulating the organism against generalization failure during subsequent extinction.
The complexity and entropy of the schedule also play a decisive role in governing associative accrual. Highly predictable, low-entropy sequences (such as alternating pairs: R-R-N-N-R-R-N-N) yield segmented aftereffect conditioning, wherein internal cues become bound to distinct sub-patterns. High-entropy, quasi-random sequences present a wide array of permutation transitions, such as N-R, N-N-R, R-R, and R-N-R. Under high sequence entropy, the internal sensory environment on any given trial is variable, compelling the instrumental response to develop associative connections across a heterogeneous pool of composite cues. This associative dispersion accounts for the vast behavioral tenacity observed under chaotic reinforcement schedules, cementing Capaldi’s assertion that sequence architecture is a prime parameter of learning.
5. Temporal Dynamics: Intertrial Intervals and Trace Decay
5.1 Intertrial Interval (ITI) as an Experimental Variable
The role of the intertrial interval (ITI) represents a critical proving ground for Capaldi’s sequential theory. If the internal aftereffect of nonreinforcement ($S^N$) is a physical stimulus trace that dissipates over time, then the duration of the ITI should theoretically dictate the survival, intensity, and conditionability of that trace. In early runway experimentation, the ITI was typically restricted to brief spans—ranging from 10 to 60 seconds—under which lingering sensory traces were readily presumed to survive. Under these massed conditions, sequential effects, including patterned running and PREE, emerged with exceptional consistency.
However, methodological critics soon challenged sequential theory by extending the intertrial interval into protracted domains of several minutes, hours, or even a full 24 hours. Classic learning models predicted that if trials were spaced by 24 hours, any putative physical sensory trace generated in the goal box would completely dissipate long before the animal was placed into the start box for the subsequent trial. Initial empirical tests appeared to validate these critiques: when ITIs were substantially widened, the magnitude of the PREE was frequently attenuated, and single-alternation patterning seemed to collapse. Opponents of sequential theory declared that while trace mechanics might operate under narrow, massed conditions, they could not explain instrumental persistence across real-world temporal intervals.
Capaldi responded with rigorous experimental programs designed to isolate the ITI from confounding variables, such as inter-trial environmental interference, differential handling, and contextual drift. He demonstrated that sequential effects could, in fact, survive protracted ITIs if appropriate structural controls were maintained. By demonstrating that animals could learn single-alternation patterning across multi-minute intervals, Capaldi showed that sequential effects were not restricted to fleeting peripheral aftereffects. The survival of sequential control across extended intervals necessitated a profound evolution in Capaldi’s conceptualization of the underlying trace mechanism, initiating a shift from passive physical decay to cognitive memory processing.
5.2 Trace Decay versus Memory Retrieval Mechanisms
To reconcile sequential theory with the reality of long-interval persistence, Capaldi elevated his model from a simple sensory-trace formulation to a sophisticated cue-dependent memory retrieval hypothesis. He posited that the physical outcome of nonreinforcement does two things: it generates an immediate short-term aftereffect ($S^N$), and it deposits a lasting memory representation into long-term storage. When an extended ITI elapses, the immediate short-term physical trace decays toward baseline; however, the internal memory representation remains intact. Capaldi redesignated this retrieved memory trace as an internal memory stimulus: $S^{Mem}$ or modified internal memory cues.
Under this expanded architecture, when an animal is returned to the apparatus after a protracted interval (even 24 hours), the static contextual apparatus cues ($S_{app}$)—including the distinctive visual, tactile, and spatial properties of the start box—act as powerful retrieval cues that reactivate the stored memory representation of the previous trial’s outcome. Upon reactivation, the retrieved memory of nonreward functions within the central nervous system as an internal stimulus cue identical in operational properties to the primary aftereffect $S^N$. Capaldi termed this process contextual reinstatement: the physical apparatus re-evokes the animal’s internal outcome memory, placing the animal back into an internal sensory state that drives conditioned approach behavior.
This theoretical advance neatly resolved the long-interval PREE anomaly. Resistance to extinction at 24-hour ITIs was not driven by a miraculous 24-hour physical sensory persistence; it was driven by the contextual retrieval of sequentially conditioned memories. When a partially reinforced animal arrives in the runway during long-interval extinction, the apparatus cues reactivate the long-term memory of prior nonreinforcements. Because those memory states were repeatedly paired with reward during acquisition (via retrieved-memory N-R transitions), the animal executes the instrumental response. Capaldi thus successfully translated a physiological sensory construct into a rigorous cognitive-mnemonic framework, maintaining behavioral objectivity while expanding theoretical reach.
5.3 Retention and Consolidation of Sequential Information
The operational durability of sequentially conditioned cues was further demonstrated through empirical investigations of retention intervals and retroactive interference. Capaldi and his students demonstrated that the associative bonds forged between $S^N$ (or its retrieved counterpart) and the approach response ($R_{app}$) possess remarkable temporal stability. Animals trained on specific sequence architectures could be removed from the experimental setting for weeks or months; upon re-testing, their extinction resistance or patterned responding remained robust, indicating that sequential learning engages stable consolidation pathways.
Crucially, sequential theory generated precise predictions regarding retroactive interference. If an animal experiences an extraneous, non-apparatus experience during the intertrial interval—such as being handled roughly, placed in an alternate holding box, or exposed to unpredicted food pellets in a different setting—the stability of the lingering sequential trace is subject to interference. Capaldi demonstrated that if an interpolated event disrupts the internal representation of the preceding trial outcome, the functional N-R transition is fractured. The animal fails to associate the subsequent runway reward with the genuine aftereffect of the alley’s nonreward, thereby attenuating both patterned running and subsequent extinction persistence.
This vulnerability to specific forms of interference confirmed that sequential traces behave according to the established laws of human and animal memory systems. The sequential aftereffect was revealed to be an active, dynamic mnemonic code. While short-term aftereffects could be overridden by intense immediate sensorimotor distractors, long-term sequential memories, once consolidated, exhibited profound resistance to natural decay. Capaldi showed that animals maintain an exquisite internal record of reinforcement history, updating this internal register dynamically trial-by-trial to navigate an uncertain world.
6. Sequential Theory versus Amsel’s Frustration Theory
6.1 Theoretical Framework of Abram Amsel’s Frustration Theory
Throughout the latter half of the twentieth century, Capaldi’s primary theoretical competitor was Abram Amsel, whose fractional anticipatory frustration theory dominated accounts of the Partial Reinforcement Extinction Effect. Amsel grounded his model in the motivational and emotional mechanics of Kenneth Spence and Clark Hull. He posited that the unexpected omission of an anticipated reward induces an unconditioned, aversive emotional reaction termed primary frustration ($R_F$). Primary frustration is characterized as an intensely visceral state possessing drive properties that motivate unconditioned avoidance, escape, or behavioral disruption.
According to Amsel, when an organism experiences intermittent reinforcement, the continuous pairing of environmental apparatus cues with primary frustration causes the animal to develop a conditioned fractional anticipatory frustration response ($r_F$). This internal anticipatory emotional state produces its own distinctive interoceptive sensory feedback, designated as $s_F$ (the frustration stimulus). Initially, $s_F$ acts as an inhibitory or repulsive signal, causing the animal to hesitate, vacillate, or attempt escape from the alleyway.
The cornerstone of Amsel’s explanation for PREE is the mechanism of counterconditioning. Under an intermittent reinforcement schedule, an animal experiencing the internal visceral sensations of frustration ($s_F$) is nevertheless forced by experimental constraints to traverse the runway, where it eventually finds food. Through repeated pairings, the instrumental approach response becomes conditioned directly to the internal frustration cue: $s_F \rightarrow R_{app}$. When extinction begins and all food is withdrawn, continuous reinforcement animals succumb to explosive frustration that halts their behavior. In contrast, partially reinforced animals experience the same frustration, but for them, $s_F$ has been transformed via counterconditioning into an internal trigger to approach. Resistance to extinction, in Amsel’s view, was a triumph of learned emotional fortitude.
6.2 Critical Points of Divergence Between Capaldi and Amsel
The debate between Capaldi’s sequential theory and Amsel’s frustration theory centered on a fundamental divide: is the Partial Reinforcement Extinction Effect governed by emotional counterconditioning, or is it governed by cognitive-mnemonic stimulus encoding? While Amsel conceptualized the internal cue ($s_F$) as an inherently emotional, motivational state born of thwarted expectations, Capaldi conceptualized the internal cue ($S^N$) as an affectively neutral, informational memory trace. For Capaldi, nonreinforcement was not necessarily a catastrophic emotional shock; it was simply an environmental outcome that left an observable physical and cognitive aftereffect.
This conceptual divergence yielded markedly different empirical predictions. Frustration theory required that the animal possess an established expectancy of reward before nonreinforcement could generate primary frustration. Consequently, Amsel’s model dictated that PREE could not occur unless an animal had already experienced substantial prior reinforced trials sufficient to build an associative expectation ($r_R – s_R$). Furthermore, frustration theory asserted that if the magnitude of the reward was vanishingly small, primary frustration would be minimal or nonexistent, meaning PREE should fail to materialize under small-reward conditions.
Capaldi attacked these vulnerabilities directly. Sequential theory did not require the prior establishment of reward expectancies to generate aftereffects; an animal could register the difference between outcomes from the very beginning of training. More critically, Capaldi pointed out that Amsel’s model had to assume complex, multi-stage emotional mechanics—primary frustration, conditioned anticipatory frustration, visceral feedback stimuli, and counterconditioning—where sequential theory required only simple stimulus-trace dynamics and elementary associative conditioning. Sequential theory offered parsimony, treating the memory of nonreward with the same analytical mechanics applied to any external visual or auditory cue.
6.3 Crucial Experiments Resolving the Controversy
To evaluate these competing models, Capaldi and his collaborators conducted a sequence of crucial experiments that placed frustration theory and sequential theory in direct empirical opposition. A primary battleground involved runway experiments utilizing exceptionally low-magnitude rewards—such as tiny single food pellets or minute concentrations of liquid sucrose. Frustration theory explicitly asserted that small rewards generate negligible emotional investment; thus, the omission of a tiny reward should fail to elicit the primary frustration necessary to drive counterconditioning. Sequential theory, conversely, predicted that even tiny outcomes leave distinct internal memory traces, meaning that N-R transitions under small-reward regimes should still generate robust resistance to extinction. The empirical results overwhelmingly supported Capaldi: robust PREE was repeatedly observed under minimal-reward conditions where emotional frustration was effectively absent.
A second decisive line of evidence emerged from precise sequence-manipulation paradigms where frustration theory predicted behavioral failure, but sequential theory predicted success. Capaldi administered schedules with identical reinforcement percentages and identical theoretical opportunities for frustration, but systematically varied whether nonrewarded trials directly preceded rewarded trials. If resistance to extinction were driven simply by running in the presence of generalized frustration, the precise ordering of trials across a multi-day training phase should have exerted minor influence. Yet, as demonstrated in Capaldi’s sequence studies, when N-R transitions were structurally eliminated (e.g., placing all nonreinforced trials at the end of the day’s training block), PREE vanished. Frustration was experienced, yet persistence failed to develop because the explicit sequential link between the memory of nonreward and the execution of approach was broken.
Furthermore, evaluations of latent extinction—where animals are placed directly into an unbaited goal box without running the runway—revealed critical sequential dependencies that frustration models could not accommodate. While modern behavioral neuroscience acknowledges that intense nonreward can certainly recruit affective, amygdalar-driven emotional responses under high-incentive conditions, the historical consensus settled on a synthesis: emotional frustration may contribute to response vigor under extreme conditions, but the fundamental architecture of the Partial Reinforcement Extinction Effect is dictated by sequential memory mechanics. Capaldi’s cognitive-mnemonic aftereffects provided the foundational associative framework that frustration theory lacked.
7. Magnitude of Reinforcement and Sequential Contrast
7.1 Interaction Between Reward Magnitude and Schedule
The magnitude of reinforcement exerts an intriguing, paradoxical influence on instrumental performance, depending fundamentally on the schedule of reinforcement. Under continuous reinforcement schedules, increasing the magnitude of the reward (e.g., from one food pellet to twenty food pellets) produces an unexpected outcome during extinction: larger rewards lead to faster behavioral extinction. This phenomenon, known as the Magnitude of Reinforcement Extinction Effect (MREE), presented another puzzle for classical learning theory, which assumed that larger rewards should deposit greater habit strength and, consequently, greater resistance to extinction.
Under partial reinforcement, however, Capaldi revealed that this dynamic inverts. Elevating the magnitude of reinforcement within an intermittent schedule does not accelerate extinction; rather, it markedly enhances the Partial Reinforcement Extinction Effect. PRF animals trained with massive rewards display phenomenal, enduring persistence that vastly outstrips PRF animals trained with modest rewards. Capaldi’s sequential theory provided a unified, non-paradoxical account for both phenomena by analyzing the cue properties of the internal aftereffects generated by varying reward magnitudes.
According to Capaldi, the transition from an outcome of large magnitude to an outcome of zero magnitude represents a massive sensory disparity. In continuous reinforcement, an animal trained with large rewards develops an internal conditioning complex heavily reliant on the intense rewarding aftereffect: $(S_{app} + S^{R-Large})$. When nonreinforcement is introduced in extinction, the resulting nonreward trace ($S^N$) represents a radical departure from the animal’s baseline internal state. The resulting generalization decrement is immediate, profound, and devastating, causing the animal to halt running almost instantly. Under partial reinforcement, however, that very same massive sensory disparity works to the animal’s advantage. An N-R transition in a large-reward schedule pairs an exceptionally salient nonreward trace ($S^N$) with an exceptionally potent reward, forging an unbreakable associative bond ($S^N \rightarrow R_{app}$). Sequential theory proved that reward magnitude scales the distinctiveness and conditionability of the internal transitions.
7.2 Simultaneous and Successive Contrast Effects
Sequential theory provided vital explanatory leverage for behavioral contrast phenomena, famously documented by Leo Crespi in 1942. In successive negative contrast (the classic Crespi effect), an animal trained to run for a large reward is suddenly downshifted to a small reward. Rather than adjusting smoothly to the running speed of a control group trained continuously on small rewards, the downshifted animal displays a dramatic behavioral plunge: its running speeds drop significantly below those of the small-reward control group. Conversely, in positive behavioral contrast, upshifting an animal from a small to a large reward yields running speeds that temporarily exceed those of animals trained exclusively on the large reward.
Traditional accounts viewed negative contrast as an acute manifestation of emotional depression or depressive frustration. Capaldi, however, demonstrated that contrast effects could be rigorously modeled as sequential aftereffect mismatch and generalization decrement. An animal running down an alleyway does so within a specific internal stimulus framework built up by its preceding reward history. When an animal accustomed to large rewards suddenly encounters a small reward, its subsequent trials do not occur in a neutral state; they occur in the presence of an unexpected, novel aftereffect state. The small reward fails to match the retrieved internal expectancy of the large reward, provoking an immediate behavioral hesitation that mirrors standard generalization decrement.
Capaldi showed that the magnitude of negative contrast is determined by the sequencing of preceding trials. If an animal is trained on an alternating schedule containing both large and small rewards, negative contrast fails to emerge when the animal is shifted permanently to small rewards. The animal has already conditioned its approach response to the aftereffects of both outcomes. Capaldi formulated algebraic baseline reference models, showing that organisms calculate an active internal standard of comparison based on sequentially experienced outcomes. Contrast is not a magical emotional overreaction; it is the observable consequence of an internal stimulus mismatch generated by the serial architecture of reinforcement.
7.3 Multiple Outcome Magnitudes within Single Sequences
To explore the limits of sequential stimulus control, Capaldi and his students engineered intricate experimental sequences incorporating multiple distinct reward magnitudes within a single training regimen. A common protocol utilized three distinct outcomes: Large reward (L), Small reward (S), and Nonreinforcement (N). By manipulating the precise serial order of these outcomes—such as L-S-N, N-S-L, or S-L-N—Capaldi demonstrated that animals do not merge these experiences into a blurry, homogenized internal average. Instead, each outcome generates a distinct, identifiable aftereffect: $S^L$, $S^S$, and $S^N$.
In schedules where specific outcomes reliably predicted subsequent events, animals exhibited sophisticated multi-level discriminative running. For instance, in an L-S-N sequence repeated across days, rats learned to run rapidly on trials following L (which predicted S), run moderately on trials following S (which predicted N), and run with extreme caution or hesitation on trials where N predicted another nonreinforcement. The intermediate aftereffect ($S^S$) was shown to possess distinct perceptual properties that occupied a verifiable position on an internal psychological continuum between $S^L$ and $S^N$.
These empirical tests confirmed the algebraic additivity and distinct identity of sequential traces. Internal states do not act as broad, undifferentiated emotional washes; they function as a refined palette of interoceptive cues. When an animal transitions across varied outcome magnitudes, the brain preserves a high-fidelity record of the preceding trial’s qualitative value. Capaldi established that animals dynamically organize these diverse aftereffects into an internal hierarchical code, proving that complex instrumental behavior is governed by precise sequences of internal sensory information.
8. Methodological Paradigms in Capaldi’s Experimental Program
8.1 The Straight-Alley Runway Paradigm
The definitive empirical vehicle for Capaldi’s experimental program was the straight-alley runway. While seemingly simple compared to modern computerized apparatuses, the straight-alley runway represented a triumph of experimental isolation and mechanical precision. The apparatus typically measured between 6 to 12 feet in length and was segmented into three distinct, functionally isolated compartments: a start box, a main running section, and a goal box. Each compartment was separated by guillotine doors that could be raised and lowered silently via mechanical linkages or automated pneumatic systems.
A primary innovation of Capaldi’s runway was the integration of automated photobeam timing systems. As the animal moved down the alley, it sequentially broke infrared photobeams situated at strategic intervals: immediately upon leaving the start box, across the middle running section, and at the entrance to the goal box. This arrangement permitted the independent, millisecond-accurate recording of three discrete behavioral latencies:
- Starting Speed: The latency between the raising of the start-box door and the interruption of the first photobeam, indexing the animal’s immediate motivation and internal cue control.
- Running Speed: The velocity through the mid-section of the alley, reflecting pure instrumental approach vigor.
- Deceleration/Goal Speed: The speed across the final segment immediately prior to goal-box entry, capturing anticipatory discrimination, hesitation, or reward expectancy.
By measuring these discrete temporal segments, Capaldi could isolate subtle sequential effects that would be obliterated in free-operant setups. If an animal was hesitating because it anticipated nonreinforcement, that hesitation appeared predominantly in the starting speed or deceleration speed. Furthermore, the discrete-trial runway completely eliminated the confounding variable of operant baseline rates; the experimenter held total authority over when a trial began, how long the animal remained in the presence of the outcome, and how much time elapsed before the next trial began. This rigorous standardization of handling, food deprivation regimens, and feeding schedules transformed the runway into a precision instrument for measuring the internal dynamics of animal memory.
8.2 Operant Free-Operant versus Discrete-Trial Comparisons
The dominance of B.F. Skinner’s free-operant methodology during the 1950s and 1960s posed a theoretical challenge: could Capaldi’s sequential theory, forged in the discrete-trial runway, apply to the continuous lever-pressing of the operant chamber? In a standard Skinner box under a variable-interval or variable-ratio schedule, an animal presses a lever continuously, emitting dozens of responses per minute. Under such free-operant conditions, identifying where one “trial” ends and another begins becomes theoretically ambiguous, and the intertrial interval collapses into a fluid, self-paced continuum.
To resolve this tension, Capaldi and his contemporaries translated runway sequential parameters into discrete-trial operant setups. In these paradigms, the lever was not continuously available; instead, it was inserted into the chamber via a mechanical solenoid to initiate a discrete trial. The animal had a set window to depress the lever once. Upon response execution (or failure), the consequence (reinforcement or nonreinforcement) was delivered, the lever was retracted, and an experimenter-controlled ITI commenced. Under these controlled conditions, all the hallmark sequential phenomena observed in the runway—patterned running (alternating fast and slow lever-press latencies), N-length effects, and the precise dependence of PREE on N-R transitions—were replicated with absolute fidelity.
These findings revealed that sequential theory was not an idiosyncratic artifact of runway locomotion. The underlying associative and mnemonic mechanics operate with equal validity across disparate motor topographies, from the whole-body locomotion of running down an alley to the focal motor execution of depressing a lever. Capaldi demonstrated that the critical determinant of instrumental behavior is not the physical motor modality, but the serial information structure governing the relationship between internal outcome memories and subsequent response opportunities.
8.3 Control Procedures and Confound Mitigation
A major threat to behavioral experimentation in rodent models is the pervasive influence of confounding sensory artifacts, most notably olfactory cues. During the late 1960s and 1970s, rigorous researchers discovered that rats running in straight-alley runways frequently leave behind distinct odor trails. An animal that encounters nonreinforcement in the goal box may deposit an odor of nonreward or frustration (e.g., via alarm pheromones or specialized footprint secretions), while a rewarded animal leaves a rewarding or neutral odor trail. Subsequent animals placed in the alley can utilize these external olfactory trails as discriminative cues, running fast over “reward odors” and slowing down over “nonreward odors,” thereby creating a powerful experimental artifact that mimics genuine sequential learning.
Capaldi recognized this threat early and instituted rigorous control procedures that became the gold standard of behavioral testing. To neutralize olfactory artifacts, Capaldi employed meticulous sanitization protocols, exhaust fans creating continuous laminar airflow through the alley, and sophisticated group-running designs. Rather than testing a single animal through an entire sequence, animals were run in squads using staggered, counterbalanced sequences. In these designs, an animal encountering a nonreinforced trial would be placed into an alley that had just been traversed by an animal experiencing a rewarded trial, systematically decoupling external physical odors from internal reinforcement history.
Additionally, Capaldi standardized goal-box confinement durations. A pervasive confound in nonreward research is that animals left in an empty goal box for differing durations experience differing degrees of temporal delay. Capaldi established strict confinement protocols—typically fixing goal-box confinement at exactly 15 or 30 seconds for both reinforced and nonreinforced trials—ensuring that total environmental exposure remained constant. Advanced repeated-measures statistical designs were developed to track trial-by-trial sequential dependencies, ensuring that observed running latencies were driven solely by the experimental manipulation of sequential outcome history rather than apparatus artifacts.
9. Mathematical Formulations and Formal Models of Sequential Theory
9.1 Quantification of Aftereffect Strength
To elevate sequential theory beyond qualitative description, Capaldi sought to formalize the dynamics of aftereffect generation, decay, and associative accrual through mathematical equations. The primary challenge lay in modeling how the internal stimulus trace of nonreinforcement ($S^N$) attenuates across time, and how successive nonreinforcements compound into higher-order traces. Capaldi modeled the effective strength of an aftereffect at the initiation of a subsequent trial as an exponential decay function governed by temporal intervals and trial parameters.
Let $S^N_0$ denote the maximal aftereffect intensity generated immediately upon goal-box exit following nonreinforcement. The intensity of the trace after an intertrial interval $t$ is expressed as:
$$S^N(t) = S^N_0 \cdot e^{-\lambda t}$$
where $lambda$ represents an empirical decay coefficient capturing species-specific decay rates, contextual stability, and retroactive interference. When a sequence presents a run of consecutive nonreinforced trials ($N_1, N_2, dots, N_k$), the internal trace compounds. The composite internal stimulus state following a run of length $k$, designated as $S^{Nk}$, is modeled as a vector summing the residual decaying traces of prior nonreinforcements plus the immediate impact of the most recent outcome:
$$S^{Nk}(t) = \sum_{j=1}^{k} S^N_{0, j} \cdot e^{-\lambda \cdot t_j}$$
where $t_j$ represents the cumulative temporal distance from the termination of trial $j$ to the current response initiation. This vector representation allows sequential theory to formally define internal states not as static binary switches, but as continuous mathematical coordinates in an internal multidimensional stimulus space.
9.2 Predictive Equations for Extinction Persistence
The core mathematical objective of sequential theory was deriving quantitative predictions for extinction persistence. Rather than relying on vague notions of “habit durability,” Capaldi formulated an index of resistance to extinction based on the cumulative associative value accrued by sequential cues during the acquisition phase. In sequential theory, the associative strength ($V$) of the instrumental approach response is partitioned between the static apparatus cues ($S_{app}$) and the dynamic internal traces ($S^N$ and $S^R$):
$$V_{Total} = V(S_{app}) + V(S^{Trace})$$
Extinction resistance is modeled as a direct function of the associative strength specifically bound to the nonreward aftereffect traces, $V(S^N)$, relative to the maximum N-length experienced during acquisition. Capaldi established that the probability of response persistence on any given extinction trial $E_m$ is governed by the associative proximity between the prevailing extinction trace $S^{N(ext)}$ and the training traces that were directly paired with reward during N-R transitions. The running speed on extinction trial $m$ ($R_m$) is modeled as:
$$R_m = \left( V(S_{app}) + \sum_{k=1}^{n} w_k \cdot V(S^{Nk}) \cdot G(S^{N(ext)}, S^{Nk}) \right) \cdot \Phi$$
where $w_k$ represents the weighting factor corresponding to the frequency of occurrence of each N-length during acquisition, $G(S^{N(ext)}, S^{Nk})$ is a Gaussian generalization function measuring the psychophysical similarity between the current extinction trace and the trained traces, and $Phi$ is an empirical scaling constant transforming net associative strength into physical running velocity.
This formulation generated predictive precision that anticipated and challenged the computational models of Robert Rescorla and Allan Wagner (1972). While the Rescorla-Wagner model operated on trial-level error-correction mechanisms, it initially struggled with sequential dependencies and micro-structural trial order because it treated trials as exchangeable statistical events within a compound conditioning block. Capaldi’s algebraic formulations explicitly accounted for the local sequence architecture, providing mathematical derivations of single-alternation discrimination ratios that classical linear-operator associative models could not capture.
9.3 Formalization of the Generalization Decrement Index
A crucial mathematical parameter in sequential theory is the Generalization Decrement Index ($GDI$), which quantifies the perceptual and associative distance between the stimulus complex encountered during acquisition and that encountered during extinction. Capaldi formalized this transition to explain why continuous reinforcement groups experience rapid extinction collapse while partial reinforcement groups persist.
Let the total internal and external stimulus complex during acquisition be designated as vector $\mathbf{S}_{acq}$, and the stimulus complex during extinction be designated as vector $\mathbf{S}_{ext}$. In a continuous reinforcement schedule, the internal component of $\mathbf{S}_{acq}$ consists purely of $S^R$ traces, whereas in extinction, the internal component shifts entirely to $S^N$. Capaldi defined the Generalization Decrement Index as an inverse function of the inner product (or cosine similarity) between the normalized acquisition and extinction stimulus vectors:
$$GDI = 1 – \frac{\mathbf{S}_{acq} \cdot \mathbf{S}_{ext}}{|\mathbf{S}_{acq}| |\mathbf{S}_{ext}|}$$
For a continuously reinforced subject, because $S^R$ and $S^N$ occupy orthogonal or distant positions within the internal sensory space, the inner product is minimal, causing $GDI$ to approach its theoretical maximum of 1.0. This extreme generalization decrement directly suppresses the behavioral output:
$$R_{ext} = R_{asy\mp} \cdot (1 – GDI)$$
As $GDI \rightarrow 1.0$, running speed $R_{ext}$ immediately plummets to zero. Conversely, for a partially reinforced subject trained on diverse N-lengths, the training stimulus vector $\mathbf{S}_{acq}$ contains dense representations of both $S^R$ and various states of $S^N$. When shifted to extinction, the extinction vector $\mathbf{S}_{ext}$ remains highly collinear with the training vector. The inner product remains large, $GDI$ approaches zero, and running velocity $R_{ext}$ persists with negligible suppression. Through these mathematical models, Capaldi converted classical, descriptive stimulus-generalization concepts into an exact, predictive quantitative science.
10. Cognitive Extensions: Memory, Expectancy, and Counting
10.1 The Transition from S-R Traces to Cognitive Representations
As sequential theory matured throughout the 1970s and 1980s, Capaldi instituted a profound theoretical transformation: the systematic re-conceptualization of the aftereffect ($S^N$) from a passive, mechanical stimulus trace into an active, cognitive internal memory representation. While his early work maintained strict neobehaviorist terminology to avoid the philosophical pitfalls of subjective mentalism, the accumulating empirical evidence demanded a richer cognitive framework. Capaldi recognized that the internal state generated by nonreinforcement was not merely an automatic physiological residual, but a functional memory code containing rich informational content regarding past events.
This theoretical evolution led to the formalization of the retrieved-memory hypothesis. Capaldi asserted that organisms do not navigate the world solely through immediate retrospective sensory feedback; they utilize retrieved memories to forge prospective outcome expectancies. An animal at the choice point or in the start box is not merely being pushed forward by a decaying physiological aftereffect; the static cues of the apparatus trigger the retrieval of an internal memory ($S^{Mem}$) that allows the organism to anticipate what outcome is likely to occur at the end of the runway.
By defining internal memories as objective, functional stimulus inputs capable of entering into associative bonds, Capaldi forged a crucial bridge between strict associative behaviorism and emerging cognitive psychology. He proved that one could investigate complex cognitive representations—such as memories, expectations, and internal reference points—without abandoning the rigorous methodological and associative laws established by classical conditioning. Sequential theory demonstrated that internal cognitive states obey the very same associative dynamics as external environmental stimuli, anchoring animal cognition within an empirical, testable framework.
10.2 Numerical Competence and Sequential Counting in Animals
One of the most remarkable extensions of Capaldi’s sequential paradigm was the demonstration of rudimentary numerical competence and sequential counting in rodent models. In a series of groundbreaking experiments conducted with his students, Capaldi investigated whether rats could track serial order and ordinal positions within complex, structured trial sequences. Rather than presenting simple alternating schedules, Capaldi engineered repetitive serial patterns composed of varying run lengths, such as R-R-R-N or R-R-R-R-N.
In a classic paradigm utilizing an R-R-R-N sequence repeated across days, rats were administered three consecutive rewarded trials followed consistently by a single nonreinforced trial. To perform optimally, the animal needed to run rapidly across Trials 1, 2, and 3, but decelerate or hesitate specifically on Trial 4. Because all external apparatus cues, handling procedures, and intertrial intervals were held perfectly identical across all four trials, the physical environment provided zero external discriminative information. The animal could rely solely on an internal serial tally of the preceding trial outcomes.
The results were unequivocal: after systematic training, rats displayed sharp, discriminative running latencies. They ran with high velocity on Trials 1, 2, and 3, but exhibited an abrupt, statistically robust deceleration specifically on Trial 4. Capaldi demonstrated that the animals were not relying on internal temporal clocks or cumulative time intervals; if the ITI between trials was randomly varied, the discriminative performance on Trial 4 remained intact. The animals were engaging in functional serial enumeration: the internal memory of having received reward once ($S^{R1}$), twice ($S^{R2}$), or three times ($S^{R3}$) served as an internal counting register. Trial 4 was identified by the presence of the internal compound memory cue $S^{R3}$, which signaled that nonreinforcement was imminent.
Furthermore, Capaldi demonstrated that animals exhibit sophisticated chunking behaviors in complex sequences. When faced with lengthy, multifaceted reward-nonreward sequences (e.g., R-R-N-N-R-R-N-N), rats organized the trial elements into structured cognitive chunks, using the transition points between chunks as primary cognitive anchors. This empirical demonstration of numerical cognition and serial list learning firmly rooted rudimentary mathematical competence within the associative mechanisms of sequential memory.
10.3 Hierarchical Organization of Sequential Memory
Capaldi’s later investigations proved that sequential memory is not organized merely as a linear, bead-on-a-string chain of isolated traces; rather, it is organized into structured, rule-governed mental representations. Animals trained on complex sequential patterns demonstrated the ability to extract overarching structural rules that could transfer across novel task domains. For example, if an animal learned an abstract monotonic rule—such as an escalating reward sequence (Small $\rightarrow$ Medium $\rightarrow$ Large) or a descending reward sequence (Large $\rightarrow$ Medium $\rightarrow$ Small)—it exhibited immediate, positive transfer when shifted to entirely novel physical environments or novel reward substances.
This capacity for structural rule learning revealed that sequential representations possess hierarchical organization. At the lowest level of the hierarchy are the specific sensory attributes of the immediate outcome (e.g., the taste, texture, and size of a particular food pellet). At the intermediate level are the relational transitions between adjacent trials (e.g., whether outcome $t$ was larger or smaller than outcome $t-1$). At the highest level is the overarching rule structure governing the entire serial episode (e.g., “continuous decrement toward nonreward”).
Capaldi and his colleagues explored the capacity limits of rodent sequential working memory, mapping the boundaries of how many sequential items an animal can hold and manipulate in an active operational state. While memory for linear strings of nonreinforced trials typically showed capacity limits around 5 to 7 discrete items (reminiscent of classical working-memory capacity limits across species), the introduction of hierarchical organization or predictable rule chunking dramatically expanded this functional capacity. These findings positioned Capaldi’s sequential paradigm as a direct precursor to modern cognitive investigations of episodic-like memory, temporal order processing, and executive control in non-human animals.
11. Neurobiological Substrates of Sequential Conditioning
11.1 Hippocampal Involvement in Trace Retention and Sequence Learning
The neurobiological validation of Capaldi’s sequential theory centered primarily on the anatomical structures mediating temporal discontiguity and relational memory, with the hippocampus occupying the absolute center of investigation. Classical associative conditioning across contiguous temporal events can proceed through subcortical and cerebellar circuits; however, sequential learning requires the preservation and active manipulation of internal stimulus aftereffects across temporal gaps (the intertrial interval). Consequently, behavioral neuroscientists recognized that the hippocampus represents the essential neural substrate for Capaldi’s internal trace mechanics.
Empirical investigations utilizing selective hippocampal lesions confirmed this dependency with extraordinary fidelity. Bilateral excitotoxic lesions of the dorsal hippocampus or complete aspiration of the hippocampal formation produce a devastating, highly specific behavioral deficit: the complete abolishment of single-alternation patterning (R-N-R-N). While hippocampally lesioned rats can acquire baseline instrumental running down a runway for continuous rewards, they are profoundly unable to learn patterned deceleration on N trials and acceleration on R trials. The lesion does not destroy the animal’s ability to run or consume reward; it destroys the capacity to maintain the aftereffect trace ($S^N$ or $S^R$) across the intertrial interval.
Furthermore, hippocampal lesions exert profound effects on the Partial Reinforcement Extinction Effect. When training occurs with distributed trials or extended intertrial intervals, hippocampally lesioned animals show a catastrophic reduction or complete absence of the PREE. At the cellular level, this deficit corresponds to the disruption of synaptic plasticity mechanisms—specifically long-term potentiation (LTP) within the CA1 and CA3 subfields. Neurophysiological recordings have revealed that CA1 and CA3 pyramidal neurons function as “time cells” and sequential trajectory encoders, firing at specific temporal coordinates within a trial gap. When the hippocampus is compromised, the temporal sequence disambiguation necessary to differentiate $S^{N1}$ from $S^{N2}$ collapses, confirming Capaldi’s postulate that sequential learning relies upon an intact, specialized temporal-mnemonic architecture.
11.2 Prefrontal Cortical Circuits and Working Memory of Outcomes
While the hippocampus provides the temporal bridging and episodic encoding of trial events, the executive maintenance and operational manipulation of sequential outcome rules require intact prefrontal cortical circuits. In the rodent brain, the medial prefrontal cortex (mPFC)—comprising the prelimbic (PL) and infralimbic (IL) cortices—works in tight reciprocal coordination with the dorsal hippocampus and the ventral tegmental area to govern sequential performance.
The prelimbic cortex is critically implicated in the active, working-memory maintenance of outcome history. Single-unit recording studies in rodents executing sequential runway and operant tasks reveal that populations of prelimbic neurons display sustained firing across the intertrial interval specifically following nonreinforced trials. This sustained neuronal activity provides the literal neurophysiological instantiation of Capaldi’s $S^N$ trace. When an unexpected nonreinforcement occurs, the sudden omission of reward triggers a shift in mPFC network dynamics, maintaining a persistent internal representation of the nonreward event that directly modulates motor output on the subsequent trial.
Conversely, the infralimbic cortex plays an indispensable role during the extinction phase itself, mediating top-down inhibitory control over conditioned responding. In Capaldi’s framework, extinction resistance requires the animal to execute an approach response in the presence of $S^N$, overriding the unconditioned competing tendencies provoked by nonreinforcement. The mPFC mediates this cognitive control by regulating downstream striatal targets, including the nucleus accumbens core. Lesions or pharmacological inactivation of the prelimbic cortex shatter the animal’s ability to track serial order and count trials in an R-R-R-N paradigm, demonstrating that the prefrontal cortex serves as the executive computational hub that tracks serial outcome vectors and executes sequential rules.
11.3 Neurochemical Modulations: Dopamine, Serotonin, and Acetylcholine
The operational mechanics of Capaldi’s sequential aftereffects are fundamentally modulated by monoaminergic and cholinergic neuromodulatory systems. Primary among these is the ascending mesocorticolimbic dopamine projection from the ventral tegmental area (VTA) to the nucleus accumbens and prefrontal cortex. The unexpected delivery of a reward generates a phasic burst in dopaminergic firing, providing the neurochemical foundation for the rewarding aftereffect $S^R$. Crucially, the unexpected omission of an anticipated reward triggers a profound, transient depression—a “phasic dopamine dip”—below baseline tonic firing rates. This dopamine dip signals negative reward prediction errors, providing an immediate, highly distinct neurochemical signature that flags the occurrence of nonreinforcement and initiates the neural cascade constituting $S^N$.
While dopamine governs prediction errors and incentive salience, the serotonergic system (originating in the dorsal raphe nucleus) acts as a critical modulator of behavioral persistence, waiting capacity, and sequential frustration tolerance. Serotonin depletion in the rodent forebrain leads to marked behavioral impulsivity and an acute vulnerability to nonreward: animals lacking normal serotonergic tone display a collapse in their ability to endure extended N-lengths, halting rapidly during extinction. Serotonin facilitates instrumental persistence by dampening the disruptive motoric reactions that initially accompany nonreward, allowing the associative bond between $S^N$ and the approach response to govern behavior.
Finally, acetylcholine (ACh) dynamics within the septohippocampal pathway mediate the crucial balance between the encoding of immediate aftereffects and the retrieval of remote sequential memories. High cholinergic tone in the hippocampus, driven by novel environmental exploration, optimizes the network for the fresh encoding of immediate trial outcomes, preventing interference from past experiences. Conversely, a drop in cholinergic tone shifts hippocampal circuitry toward a retrieval-dominant state, facilitating the contextual reinstatement of remote sequential memories ($S^{Mem}$) across extended intertrial intervals. Pharmacological manipulations using muscarinic receptor antagonists, such as scopolamine, reliably disrupt single-alternation patterning and eliminate the PREE under long-interval conditions, providing neurochemical corroboration for Capaldi’s dual-process trace encoding and memory retrieval architecture.
12. Legacy, Critiques, and Modern Applications of Sequential Theory
12.1 Historical Critiques and Theoretical Boundaries
Despite its theoretical precision and experimental triumphs, Capaldi’s sequential theory was not without vigorous opposition. Throughout the 1960s and 1970s, critics leveled challenges regarding the physiological plausibility of static stimulus traces. Early formulations of sequential theory were accused of treating the internal aftereffect as an ethereal, convenient construct that could be assigned arbitrary decay parameters post-hoc to fit any anomalous dataset. Skeptics questioned how a simple physiological trace could retain its pristine physical integrity across hours or days in the face of continuous, blooming sensory input from the external world.
A second substantial critique focused on the potential explanatory circularity of unobservable internal cues. Traditional radical behaviorists, adhering strictly to Skinnerian operant frameworks, cautioned that postulating internal aftereffects ($S^N$) risked smuggling cognitive mentalism back into behavioral science under the guise of pseudo-objective notation. If every failure of an animal to extinguish could be explained simply by asserting that $S^N$ was still present, and every extinction collapse explained by asserting that $S^N$ had decayed, the theory threatened to become unfalsifiable unless strictly constrained by mathematical parameters.
Furthermore, theoretical boundaries emerged regarding the applicability of sequential theory in complex, multi-operant ecological environments. In the real world, foraging organisms rarely encounter neat, discrete, linear runways; they navigate dynamic, patch-depletion environments characterized by concurrent, continuous behavioral decisions. Capaldi and his colleagues responded to these criticisms with an avalanche of rigorous empirical studies. They proved that the decay parameters of $S^N$ were not arbitrary, demonstrating that trace dissipation followed lawful, predictable mathematical functions across controlled retention intervals. They operationalized every sequential construct through direct physical manipulations, successfully defending the theory against charges of circularity and establishing sequential analysis as an enduring paradigm of experimental psychology.
12.2 Impact on Contemporary Reinforcement Learning and Machine Learning
In a historical realization of theoretical convergence, Capaldi’s sequential theory is now recognized as a direct foundational ancestor to modern computational reinforcement learning (RL) and algorithmic machine learning. When computer scientists Richard Sutton and Andrew Barto formulated the mathematics of Temporal Difference (TD) learning, they confronted the very same core dilemma that Capaldi had tackled decades earlier: the credit assignment problem across non-contiguous temporal states.
In modern reinforcement learning, an agent operating in an environment must learn which actions to take to maximize cumulative future rewards. When rewards are sparse, delayed, or intermittent, an agent struggles to determine which specific prior actions were responsible for the eventual outcome. To solve this, computational RL relies fundamentally on the concept of the eligibility trace—a decaying mathematical memory of past state visits and action executions:
$$e_t(s) = \gamma \lambda e_{t-1}(s) + \mathbf{1}(S_t = s)$$
The mathematical architecture of the eligibility trace is conceptually identical to Capaldi’s decaying sequential trace $S^N(t)$. In both frameworks, the system preserves an active, fading internal representation of prior non-reinforced states, allowing reward feedback received at a later point in time to backpropagate and attach directly to the lingering trace.
Moreover, Capaldi’s sequential theory laid the conceptual groundwork for solving non-Markovian decision processes in artificial intelligence. A system is non-Markovian when the current observable state is insufficient to make optimal predictions; the agent must know the historical sequence of prior states to understand the true environmental contingency. Modern deep reinforcement learning algorithms address this by incorporating recurrent neural networks (RNNs) or Transformer-based architectures that encode sequential history into an internal hidden state vector ($h_t$). This computational hidden state vector is the functional equivalent of Capaldi’s compound sequential trace ($S^{Nk}$). In engineering artificial agents capable of unprecedented persistence under extremely sparse, intermittent reward conditions, contemporary machine learning practitioners routinely deploy synthetic sequential conditioning principles pioneered by E.J. Capaldi.
12.3 Enduring Contributions to Behavioral Science
The enduring legacy of E.J. Capaldi extends far beyond the historical boundaries of the neobehaviorist runway wars. Capaldi fundamentally transformed how the behavioral sciences conceptualize reinforcement schedules. Prior to his work, experimental psychology was largely trapped in a molar paradigm, treating intermittent reinforcement as a static percentage or a broad motivational state. Capaldi permanently dismantled this oversimplification, establishing that micro-event sequencing—the trial-by-trial serial architecture of experience—is a primary experimental parameter that governs learning, memory, and performance.
Beyond schedule architecture, Capaldi performed a monumental service to theoretical psychology by seamlessly integrating associative learning mechanics with cognitive memory frameworks. Rather than viewing associative conditioning and cognitive memory as mutually exclusive paradigms, Capaldi showed that internal memories are themselves functional stimuli that obey the lawful principles of associative transfer, generalization, and extinction. He stripped the study of animal memory of ungrounded mentalism, demonstrating that the contents of an animal’s mind could be measured, tracked, and mathematically modeled through the rigorous analysis of sequential behavior.
Ultimately, Egidio John Capaldi stands as one of the definitive architects of quantitative behavioral theory. His sequential framework brought mathematical elegance, operational discipline, and explanatory unity to empirical phenomena that had bewildered researchers for decades. By proving that what an organism learns is inseparable from the chronological tapestry of its reinforcement history, Capaldi permanently altered the scientific trajectory of learning theory, leaving an indelible imprint that continues to inform comparative psychology, behavioral neuroscience, and the computational horizons of artificial intelligence.
Conclusion
The sequential theory of instrumental learning, formulated and championed across decades of relentless empirical research by E.J. Capaldi, represents one of the crowning theoretical achievements of twentieth-century behavioral science. By dismantling the classical dogma that habit strength is merely the passive, molar accumulation of reinforced responses, Capaldi redirected the analytical gaze of experimental psychology toward the micro-structural architecture of sequential experience. His foundational insight—that the absence of reward is not an empty, inactive state, but an active, stimulus-generating event that deposits persistent internal aftereffects—demolished the enduring paradox of the Partial Reinforcement Extinction Effect and provided a unified, parsimonious account of sequential patterning, contrast effects, and behavioral persistence.
Throughout its theoretical maturation, sequential theory achieved what few psychological frameworks have accomplished: it successfully bridged the ideological chasm between radical, stimulus-based behaviorism and the emerging tenets of cognitive science. Capaldi proved that internal cognitive constructs—memories, expectancies, sequential order tracking, and rudimentary numerical counting—could be investigated with absolute operational rigor, showing that internal memory states function as genuine discriminative cues governed by lawful associative and mathematical principles. As modern neuroscientists uncover the cellular time cells of the hippocampus and prefrontal working-memory networks, and as machine learning engineers refine temporal-difference algorithms and eligibility traces in artificial intelligence, the core tenets of sequential theory continue to resonate. E.J. Capaldi demonstrated that behavior is not simply a reaction to the physical present, but an exquisite, mathematically disciplined manifestation of sequential reinforcement history.
References
- Amsel, A. (1958). The role of frustrative nonreward in noncontinuous reward situations. Psychological Bulletin, 55(2), 102–119. https://doi.org/10.1037/h0043125
- Amsel, A. (1962). Frustrative nonreward in partial reinforcement and discrimination learning: Some recent history and a theoretical extension. Psychological Review, 69(4), 306–328. https://doi.org/10.1037/h0046248
- Capaldi, E. J. (1966). Partial reinforcement: A hypothesis of sequential effects. Psychological Review, 73(5), 459–477. https://doi.org/10.1037/h0023684
- Capaldi, E. J. (1967). A sequential hypothesis of instrumental learning. In K. W. Spence & J. T. Spence (Eds.), The Psychology of Learning and Motivation (Vol. 1, pp. 67–156). Academic Press. https://doi.org/10.1016/S0079-7421(08)60513-8
- Capaldi, E. J. (1971). Memory and learning: A sequential viewpoint. In W. K. Honig & P. H. R. James (Eds.), Animal Memory (pp. 111–154). Academic Press. https://doi.org/10.1016/B978-0-12-354150-5.50009-8
- Capaldi, E. J. (1993). Animal number abilities: Implications for a general model of animal approach. In S. T. Boysen & E. J. Capaldi (Eds.), The Development of Numerical Competence: Animal and Human Models (pp. 191–209). Lawrence Erlbaum Associates.
- Capaldi, E. J., & Hart, D. (1962). Influence of a previously conditioned stimulus on running speed in extinction. Journal of Experimental Psychology, 63(1), 22–26. https://doi.org/10.1037/h0044673
- Capaldi, E. J., & Miller, D. J. (1988). Counting in rats: An internal clock or an internal counter? Journal of Experimental Psychology: Animal Behavior Processes, 14(1), 3–17. https://doi.org/10.1037/0097-7403.14.1.3
- Crespi, L. P. (1942). Quantitative variation of incentive and performance in the white rat. The American Journal of Psychology, 55(4), 467–517. https://doi.org/10.2307/1417120
- Hull, C. L. (1943). Principles of Behavior: An Introduction to Behavior Theory. Appleton-Century-Crofts.
- Humphreys, L. G. (1939). The effect of random alternation of reinforcement on the acquisition and extinction of conditioned eyelid reactions. Journal of Experimental Psychology, 25(2), 141–158. https://doi.org/10.1037/h0058138
- Rawlins, J. N. P., Feldon, J., & Gray, J. A. (1980). The effects of hippocampectomy on the partial reinforcement extinction effect in rats. Experimental Brain Research, 38(3), 273–283. https://doi.org/10.1007/BF00236646
- Rescorla, R. A., & Wagner, A. R. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In A. H. Black & W. F. Prokasy (Eds.), Classical Conditioning II: Current Research and Theory (pp. 64–99). Appleton-Century-Crofts.
- Sheffield, F. D. (1949). Hilgard’s critique of Guthrie’s theory of learning. Psychological Review, 56(5), 284–291. https://doi.org/10.1037/h0054366
- Skinner, B. F. (1938). The Behavior of Organisms: An Experimental Analysis. Appleton-Century.
- Spence, K. W. (1956). Behavior Theory and Conditioning. Yale University Press. https://doi.org/10.1037/10026-000
- Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press.
- Tolman, E. C. (1948). Cognitive maps in rats and men. Psychological Review, 55(4), 189–208. https://doi.org/10.1037/h0061626