The trajectory of twentieth-century experimental psychology was defined by a profound ontological struggle regarding the nature of learning, memory, and behavioral adaptation. For decades, the dominant behavioral paradigm treated the organism as an epistemological black box. Governed by peripheral reflex arcs and operationalized through input-output pairings, animal learning was long conceptualized as the mechanical stamping-in of motor habits. Within this rigid framework, learning was asserted to consist exclusively of direct connections established between environmental inputs and muscular or glandular outputs. Internal cognitive states, internal dynamic representations, and predictive psychological expectancies were systematically dismissed as unscientific epiphenomena. This radical behaviorist consensus maintained that an animal exposed to classical conditioning acquired an unmediated reflex: the Conditioned Stimulus (CS) came to elicit the Conditioned Response (CR) via direct associative channels that completely bypassed central representations of the outcome.
The dismantling of this peripheralist dogma required methodological innovations capable of rendering internal, unobservable cognitive structures empirically undeniable. Chief among the architects of this cognitive revolution in animal learning was Robert A. Rescorla. Working at the interface of rigorous experimental methodology and emerging cognitive theory, Rescorla revolutionized associative learning by demonstrating that conditioning is not a passive mechanical registration of temporal proximity, but an active, computational encoding of statistical contingency, informational predictive value, and internal representations. Among his many experimental contributions, Rescorla’s 1973 Unconditioned Stimulus (US) devaluation experiment stands as a masterwork of theoretical isolation and behavioral architecture, providing irrefutable empirical proof that organisms encode rich, dynamic internal representations of the events they encounter.
By conceptually dissociating the learned cue from the current motivational and sensory value of its associated outcome, the US devaluation paradigm resolved a decades-old debate between Stimulus-Response (S-R) mechanistic formulations and Stimulus-Stimulus (S-S) representational architectures. Rescorla demonstrated that when an unconditioned stimulus is rendered biologically neutral or hedonically aversive through post-conditioning manipulation—in the total absence of the conditioned stimulus—the subsequent presentation of the conditioned stimulus results in an immediate, unprompted attenuation of the conditioned response. This comprehensive treatise explores the historical antecedents, mechanical nuances, computational interfaces, neurobiological substrates, and enduring clinical legacies of Rescorla’s US devaluation paradigm, illustrating how a deceptively elegant experimental protocol fundamentally remapped our understanding of the associative mind.
1. Introduction to US Devaluation and Robert Rescorla’s Paradigm
1.1 Conceptual Definition of Unconditioned Stimulus Devaluation
Unconditioned Stimulus (US) devaluation is an experimental protocol within associative learning psychology designed to probe the underlying representational structure of Pavlovian and instrumental conditioning. In standard classical conditioning, an organism is repeatedly exposed to a neutral Conditioned Stimulus (CS), such as a tone or a light, temporally paired with a biologically significant Unconditioned Stimulus (US), such as food, water, or an aversive footshock or loud noise. Over repeated pairings, the organism acquires a Conditioned Response (CR) to the CS. The foundational theoretical question that arises is: what has the organism fundamentally learned? Does the CS directly elicit the CR through an autonomous sensorimotor pipeline, or does the CS evoke a rich central representation of the US, which in turn calculates, generates, and drives the CR based on the current motivational status of that outcome?
US devaluation directly interrogates this question by selectively diminishing the affective, sensory, or motivational value of the primary reinforcer (the US) after the associative relationship between the CS and US has been thoroughly established. Critically, this post-conditioning revaluation is executed entirely in the absence of the CS. The experimenter employs secondary interventions—such as systemic illness induction via chemical agents, high-density acoustic habituation protocols, or sensory-specific satiety regimens—to alter the subjective utility of the US. If the conditioned response is sustained entirely by an autonomous S-R reflex loop, the subsequent presentation of the CS should yield an unmodified, habitual response. If, however, the CS accesses an internal representation of the US whose value has been dynamically downgraded, the response should immediately plummet without requiring any corrective learning trials.
It is vital to distinguish operational US devaluation from direct experimental extinction. In traditional extinction protocols, the conditioned stimulus is repeatedly presented in the absence of the unconditioned reinforcer (CS-alone presentations). Extinction directly exposes the organism to an overt mismatch between prediction and reality, driving new inhibitory learning that actively suppresses the original conditioned response. Conversely, US devaluation leaves the CS entirely untouched during the intermediate value-updating phase. The organism receives no direct experiential evidence that the predictive contingency between the CS and the US has altered. Therefore, any modification in the conditioned response observed during a non-reinforced probe test reflects an internal cognitive calculation—an inferential synthesis of the CS-US associative memory and the newly updated US evaluation.
By isolating the cognitive representation of the unconditioned stimulus from reflexive peripheral motor outputs, the devaluation paradigm operationalizes an analytical wedge between motor memory and outcome expectancy. It challenges the assumption that conditioned responses are homogeneous, ballistic muscular discharges. Instead, the paradigm establishes that Pavlovian behavior is dynamically computed, sensitive to homeostatic imperatives, and subservient to the real-time hedonic and survival-relevant currency of the internal representation evoked by environmental signals.
1.2 Robert Rescorla and the Cognitive Revolution in Conditioning
During the mid-twentieth century, experimental psychology was dominated by radical and neobehaviorist paradigms championed by figures such as B.F. Skinner and Clark L. Hull. These paradigms explicitly rejected internal mentalistic constructs, conceptualizing learning exclusively as changes in response probability or the accumulation of habit strength. Against this orthodox landscape, Robert A. Rescorla emerged as a transformative figure who utilized the rigorous, highly quantitative methods of animal behavior to validate the principles of the nascent cognitive revolution.
Rescorla’s fundamental epistemological critique centered on the behavioral denial of information processing. Traditional behaviorism asserted that an animal was a passive recipient of contiguous environmental stimulation, reflexively forming associative links whenever two events occurred closely together in time and space. In contrast, Rescorla posited that organisms act as rational inferential observers of environmental statistical regularities. Animals do not merely register raw physical contiguity; they process contingency, extracting correlation, predictive utility, and relative validity from complex sensory streams.
His early empirical work in the late 1960s thoroughly dismantled the temporal contiguity doctrine by demonstrating that an animal only forms an association to a conditioned stimulus if that stimulus provides genuine informational value regarding the arrival of the unconditioned event. If a shock is equally likely to occur in the presence or absence of a tone, no conditioning occurs, despite numerous instances of temporal pairing. This demonstrated that associative systems compute contingency matrices. By the early 1970s, Rescorla extended this cognitive orientation from contingency assessment to internal representational architecture, culminating in his groundbreaking 1973 devaluation experiments at Yale University.
Rescorla’s scholarship demonstrated that acknowledging mental representations does not require sacrificing experimental rigor or operational clarity. By designing elegantly controlled behavioral assays, he proved that internal cognitive representations could be mathematically, behaviorally, and mechanistically isolated, measured, and verified. His work altered the trajectory of comparative cognition, providing the operational blueprints that allowed cognitive neuroscience, modern reinforcement learning theory, and neurobiology to locate the physical substrates of expectation and mental simulation in the mammalian central nervous system.
1.3 Core Research Questions Addressed by the Paradigm
Rescorla designed the US devaluation paradigm to address fundamental questions regarding the associative architecture of the mind. At the core of his inquiry was the definitive resolution of the historical debate between Stimulus-Response (S-R) models and Stimulus-Stimulus (S-S) models of learning. Specifically, the paradigm was formulated to answer whether a conditioned stimulus forms a direct, hardwired associative bond with the downstream motor commands that manifest the conditioned response, or whether it creates an associative link to an abstract, central memory node representing the properties and value of the unconditioned stimulus.
A second vital question focused on the functional dependency of the conditioned response on real-time reinforcer value. When an animal exhibits a conditioned response, does that behavior reflect an automated, autonomous motor program that executes blindly regardless of the current biological utility of the outcome? Or is the conditioned response an emergent, dynamically calculated action that is continuously adjusted based on the current hedonic, metabolic, and motivational value assigned to the mental representation of the US? Resolving this question would establish whether associative memories are rigid, static historical records of past reinforcement, or flexible, editable cognitive systems capable of prospective integration.
Furthermore, Rescorla aimed to determine whether associative structures could be modified offline, completely detached from direct associative interaction. Could an animal update its behavioral response to a predictive cue without ever experiencing that cue in conjunction with the altered outcome? Demonstrating such capability would prove that organisms possess internal associative networks that can recombine disparate experiences inferentially. An animal could experience Event A leading to Event B, independently experience Event B acquiring negative valence, and subsequently infer—upon the isolated presentation of Event A—that Event A’s ultimate consequence is now undesirable.
Finally, the US devaluation paradigm provided an empirical diagnostic criterion to delineate the operational boundaries between automated habits and goal-directed actions. In contemporary behavioral science, this distinction forms the bedrock of our understanding of decision-making, neuroeconomics, and psychiatric pathology. Rescorla’s methodology established the empirical gold standard for testing whether an organism’s behavior is under the purposive control of an outcome representation or has devolved into an autonomous, reflexive behavioral routine.
2. Theoretical Background: S-R Versus S-S Associative Models
2.1 The Hullian Stimulus-Response Framework
The Stimulus-Response (S-R) theoretical framework reached its zenith under the mathematical systematization of Clark L. Hull. In Hull’s monumental neo-behaviorist formulation, learning is conceptualized as the gradual accretion of habit strength ($sHr$) between a sensory stimulus and an effector response. This process is driven primarily by drive reduction: whenever a motor response occurs in temporal proximity to a stimulus and is immediately followed by a biological drive reduction—such as the satiation of hunger, thirst, or escape from tissue-damaging stimulation—the connection between the sensory input and that specific motor sequence is reinforced.
Within this Hullian architecture, the conditioned stimulus acts as an autonomous mechanical trigger. Once habit strength has been stamped into the nervous system through repeated drive-reducing trials, the presence of the conditioned stimulus inevitably activates the neural circuit that drives the motor apparatus. Crucially, Hull’s model asserts that the associative bond completely bypasses higher-order central representations of the reinforcer. The unconditioned stimulus serves merely as a catalyst—a mechanistic reinforcing agent whose sole function is to cement the peripheral S-R link. Once that link is forged, the physical properties, identity, and current motivational relevance of the US are theoretically irrelevant to the elicitation of the response.
The fundamental vulnerability of strict Hullian S-R mechanisms lies in their predictive rigidity. If an associative connection is strictly stamped between a stimulus and a response, the vigor of the conditioned response should depend exclusively on the accumulated habit strength modulated by the organism’s general, non-specific physiological drive state ($D$). The model predicts that if habit strength is high, presenting the stimulus will automatically elicit the response, regardless of whether the specific outcome that originally reinforced the behavior is currently desirable. Hullian theory struggles to explain situations where an animal selectively abstains from a specific conditioned response while remaining in a state of high general arousal and general drive. Because the model lacks an internal representational node dedicated to the specific identity and qualitative dimensions of the outcome, it predicts behavioral persistence when environmental conditions demand dynamic behavioral recalibration.
2.2 Tolmanian Purposive Behaviorism and Cognitive Mapping
In stark opposition to Hull’s mechanical peripheralism stood the purposive behaviorism of Edward C. Tolman. Tolman rejected the notion that animals are passive bundles of reflexes driven mechanically through environmental mazes. Instead, he argued that organisms are active information seekers who develop rich internal representations of their environment, famously termed cognitive maps. Learning, according to Tolman, does not consist of forging direct S-R motor linkages; rather, it involves acquiring internal expectations regarding “what leads to what.”
Tolman posited that conditioning establishes internal sign-gestalt expectations. A conditioned stimulus is not a direct trigger for a muscular spasm; it is a “sign” that points to a “significate”—an internal representation of the expected outcome. Through environmental exploration and associative pairing, organisms learn that a specific environmental cue signifies the impending arrival of a specific outcome possessing definite qualitative, sensory, and affective attributes. The subsequent behavioral output is an active, purposive navigation toward or away from that expected outcome, mediated by the animal’s current motivational desires and mental representations.
Despite the conceptual appeal of Tolman’s purposive behaviorism, his framework historically suffered from a critical vulnerability: a lack of quantitative mathematical formalization and unambiguous experimental paradigms capable of ruling out sophisticated S-R counter-interpretations. Mechanistic behaviorists routinely argued that phenomena such as Tolman’s latent learning could be re-explained through fractional anticipatory goal responses ($r_g – s_g$ mechanisms)—internalized micro-proprioceptive feedback loops that preserved the core tenets of S-R theory without resorting to cognitive representations. Tolman’s assertion that an animal possesses central expectancies remained an abstract theoretical claim until researchers developed experimental designs that could empirically isolate an internal representation from peripheral motor feedback. It was precisely this operational gap that Robert Rescorla bridged by formalizing and executing the post-conditioning US devaluation experiment.
2.3 Theoretical Divergence on Associative Topology
The ideological clash between Hull and Tolman can be understood as a fundamental divergence regarding the anatomical and functional wiring diagram of the learning brain. The S-R associative topology proposes a direct, horizontal feedforward pathway: Sensory Node (CS) $\rightarrow$ Motor Control Unit (CR). The unconditioned stimulus acts externally upon this pathway, functioning as an energetic valve that potentiates the synaptic weight between sensory input and motor execution. In this model, there is no representational node representing the US within the operational circuit that generates the CR during recall; the US is merely a historical ghost that reinforced the connection.
Conversely, the S-S associative topology proposes a triangular, mediated network: Sensory Node (CS) $\rightarrow$ Internal Central Representation of the Outcome (US) $\rightarrow$ Motor/Behavioral Output Center (CR). In this structural architecture, the CS does not possess an autonomous channel directly to the motor effectors. Instead, its sole functional capability is to reactivate the central neural representation of the unconditioned stimulus. Once reactivated, this US representation assesses its own current biological valence, sensory features, and affective charge. If the US node is currently evaluated as appetitive, dangerous, or salient, it sends an activating projection to downstream behavioral execution circuits. If the US representation has been neutralized, deactivated, or inverted, the signal terminates at the representational stage, yielding no downstream behavioral expression.
These divergent topologies generate mutually exclusive, highly testable empirical predictions when subjected to post-acquisition outcome devaluation. If an animal is conditioned with a CS paired with a US, and the US is subsequently devalued in the total absence of the CS, the two models make radically opposing predictions:
- The S-R Prediction: Because the physical associative connection was stamped directly between the sensory input of the CS and the motor output of the CR, the internal motivational status of the US is entirely inconsequential. Upon testing with the CS alone, the habit strength remains fully intact. The CR must fire with unattenuated vigor.
- The S-S Prediction: Because the CS operates strictly by activating the central representation of the US, and because that representation has been independently updated to a state of zero or negative subjective value, the transmission of behavioral activation will fail. Upon testing with the CS alone, the CR must show immediate, unprompted attenuation.
The theoretical power of this paradigm rests upon testing the animal without direct CS-devalued US re-pairing. The animal must not experience the CS followed by the altered US, as that would permit standard first-order extinction or counterconditioning. The critical probe must be conducted during an extinction test, forcing the animal to rely solely on internal computational inference.
3. The Historical Context of Classical Conditioning in the 1960s and 1970s
3.1 The Dogma of Temporal Contiguity
For more than half a century following the translation of Ivan Pavlov’s foundational works on conditioned reflexes, experimental psychology rested upon the bedrock assumption of temporal contiguity. Pavlov’s experiments demonstrated that if a conditioned stimulus preceded an unconditioned stimulus by a brief temporal interval, an associative reflex was forged. This led to the pervasive scientific dogma that temporal proximity was both necessary and sufficient for associative learning to occur. If two neural events fired concurrently or in close temporal sequence, associative mechanisms were assumed to bind them together mechanically.
Throughout the 1950s and 1960s, however, anomalies accumulated that the contiguity doctrine could not comfortably accommodate. John Garcia’s discovery of Conditioned Taste Aversion demonstrated that animals could associate a gustatory CS with an emetic US across intervals spanning several hours, completely shattering the brief temporal windows prescribed by contiguity purists. Concurrently, experimentalists noted that presenting thousands of contiguous CS-US pairings did not reliably generate conditioning if the environmental context already predicted the outcome, or if another predictive cue was present.
Robert Rescorla dealt the fatal blow to the contiguity dogma in his landmark 1968 paper, “Probability of shock in the presence and absence of CS in fear conditioning.” Rescorla systematically held temporal contiguity constant while manipulating the statistical contingency between the CS and the US. He exposed groups of rats to an identical number of contiguous CS-US pairings, but systematically varied the probability of the US occurring during the inter-trial intervals (the background periods when the CS was absent). The results were definitive: when the probability of the shock occurring during the CS was identical to the probability of the shock occurring in its absence, absolutely zero conditioning occurred, despite numerous contiguous pairings. Rescorla demonstrated that animals compute statistical dependencies: associations are driven not by contiguous pairings, but by the diagnostic, informational utility of the cue.
3.2 The Emergence of the Rescorla-Wagner Model
The theoretical necessity to operationalize contingency mathematically led to the development of the Rescorla-Wagner Model of conditioning, published in 1972 in collaboration with Allan Wagner. This formal computational architecture conceptualized associative learning as a mathematical process driven by prediction error:
$$\Delta V = \alpha \beta (\lambda – \sum V)$$
In this classic formulation, $\Delta V$ represents the change in associative strength between the conditioned stimulus and the unconditioned stimulus on a given trial. The parameters $\alpha$ and $\beta$ represent the salience of the CS and the learning rate parameter of the US, respectively. The asymptote of associative strength that the US can support is designated by $lambda$, while $\sum V$ represents the aggregate associative strength of all cues present on that specific trial. The difference between the actual outcome ($lambda$) and the aggregate prediction ($\sum V$) constitutes the prediction error.
The model brilliantly formalized phenomena such as blocking, overshadowing, conditioned inhibition, and extinction, securing its place as one of the most influential formal models in the history of behavioral and cognitive science. Learning occurred only when the outcome was surprising—that is, when the organism experienced a discrepancy between what it expected and what it actually received.
However, the Rescorla-Wagner model harbored an inherent representational limitation: it was fundamentally an error-correcting, single-variable mapping algorithm. In its baseline 1972 formulation, associative strength ($V$) was treated as a scalar quantity of general associative loading. The model did not explicitly maintain a multidimensional, qualitative representation of the US within the operational parameters of $V$. The asymptote $lambda$ was treated as a fixed physical parameter dictated by the external unconditioned stimulus during actual delivery. The mathematical formulation lacked an organic computational mechanism to account for representational updates that occur offline—specifically, how a change in the internal representation of $lambda$ could instantly cascade through the network to update responding in the absence of any prediction-error-generating trials ($\Delta V = 0$). It was this algorithmic limitation that underscored the necessity of the 1973 devaluation experiments, which forced learning theory to integrate qualitative mental representations into associative computational architectures.
3.3 Institutional Reception and Paradigm Shifts
The introduction of the US devaluation paradigm in the early 1970s encountered institutional resistance from traditional behaviorist circles. Radical behaviorists trained in the operant traditions of Skinner and mechanists adhering to Hullian dogma viewed the vocabulary of “representations,” “expectancies,” and “mental nodes” as a regressive slide into mentalism. They argued that theoretical psychology was abandoning hard-won operational purity in favor of hypothetical internal homunculi.
However, Rescorla’s experimental paradigms were methodologically unassailable. He did not rely on subjective phenomenological reports or introspective human psychology; he utilized impeccably controlled behavioral assays involving non-human animals (predominantly laboratory rodents) executing precisely measured behaviors within standardized testing apparatuses. The dependent variables—licking latencies, bar-press rates, and suppression ratios—were as quantitatively rigorous as any measurement demanded by the strictest behaviorist orthodoxy. The inescapable reality of the empirical data forced a paradigm shift: one could no longer claim that operational behaviorism required the denial of internal mental structures.
As the 1970s progressed, the US devaluation paradigm catalyzed a broader re-evaluation of non-human animal cognition across the global experimental psychology community. Researchers recognized that classical conditioning was not a primitive, sub-cognitive biological artifact reserved for autonomic reflexes, but a sophisticated sensory-cognitive system designed to build an internal model of the external physical environment. The devaluation paradigm became the foundational blueprint that bridged historical Pavlovian conditioning with contemporary cognitive neuroscience, directly paving the way for modern investigations into the neural bases of model-based decision systems, neuroeconomic value assignment, and frontostriatal computational loops.
4. Rescorla’s Seminal Experimental Methodology and Design
4.1 Phase One: First-Order Pavlovian Conditioning
In his landmark 1973 study, entitled “Conditioned Inhibition and Extinction” and elaborated in subsequent investigations including his definitive 1974 experiments, Rescorla established an experimental architecture structured into three distinct, non-overlapping phases. The objective of Phase One was to establish a robust, unambiguous first-order conditioned association between an arbitrary conditioned stimulus and a biologically potent unconditioned stimulus within a tightly regulated environmental context.
To prevent sensory overlap and sensory preconditioning confounds, Rescorla carefully selected conditioned stimuli from different sensory modalities. Typically, an auditory stimulus (such as a distinct 1,800 Hz pure tone or a continuous white noise track) and a visual stimulus (such as an intermittent flashing light) were employed. These cues were rigorously counterbalanced across experimental cohorts to ensure that physical salience, inherent sensory biases, or natural behavioral prepotencies could not systematically skew the quantitative outcomes.
For the unconditioned stimulus, Rescorla utilized an intense, high-decibel acoustic event: a catastrophic loud noise delivered at approximately 100 to 120 decibels through high-fidelity overhead speakers inside an acoustic isolation chamber. This acoustic US was an exceptionally potent aversive stimulus that elicited a reflexive, unconditioned defensive startle response accompanied by a profound suppression of ongoing appetitive behavior. Unlike electrical footshocks, which involve nociceptive tissue stimulation and complex neuromuscular pathways, an intense auditory US provided a unique experimental advantage: it was a stimulus to which mammalian organisms could be thoroughly habituated via prolonged, repeated exposure without inducing peripheral physical trauma or chronic systemic stress.
Phase One conditioning took place while the subjects—typically male Sprague-Dawley rats—were engaged in a steady baseline operant behavior: licking a drinking tube containing a liquid nutrient or pressing a lever on a variable-interval (VI) schedule for food reinforcement. Rescorla administered discrete pairings of the CS (e.g., a 30-second tone) co-terminating with the delivery of the loud noise US. Training regimens were distributed across multiple daily sessions, utilizing variable inter-trial intervals (ITIs) averaging several minutes to prevent temporal conditioning to contextual cues. Conditioning continued until strict behavioral criteria were met: the presentation of the CS caused near-total cessation of baseline appetitive behavior. This suppression was mathematically calculated via the classic conditioned suppression ratio ($SR$):
$$SR = \frac{B}{A + B}$$
Here, $B$ represents the response rate (e.g., licks or lever presses) during the presence of the CS, and $A$ represents the response rate during an equivalent baseline period immediately preceding the onset of the CS. A suppression ratio of $0.50$ signifies complete absence of conditioning (no difference between baseline and CS periods), whereas a ratio approaching $0.00$ signifies absolute conditioned fear and total suppression of ongoing behavior. Rescorla trained the animals until all experimental and control groups achieved uniformly low, asymptotic suppression ratios, confirming that the CS had established maximum associative control over the defensive conditioned response.
4.2 Phase Two: Independent Post-Conditioning Devaluation
Phase Two represented the critical experimental intervention: the independent, offline devaluation of the unconditioned stimulus. Having forged a powerful associative connection between the CS and the loud noise US, Rescorla separated the experimental cohort into carefully matched groups to isolate the variable of US value.
The experimental devaluation group was subjected to an extensive acoustic habituation regimen designed to systematically exhaust the biological and psychological salience of the loud noise US. To avoid contextual contamination, this habituation was conducted either in an entirely different sensory environment or within the standard chambers under altered environmental parameters. The experimental animals were presented with hundreds of massive, repeated, standalone exposures to the identical loud noise US used in Phase One. Critically, the conditioned stimulus was entirely absent throughout this entire phase; the experimental animals did not hear a single tone or see a single flashing light during Phase Two. The noise was presented at high frequencies until the unconditioned startle response and the unconditioned suppression elicited by the standalone US were completely extinguished. The animal was systematically adapted until the loud noise had been reduced from a terrifying, defensive-eliciting event to a biologically neutral, ignorable auditory artifact.
Simultaneously, control cohorts were maintained under rigorous conditions to account for potential confounding variables:
- Non-Devalued Control Group: Animals were placed in identical experimental chambers for equivalent durations but received zero exposures to the loud noise US, preserving the unconditioned stimulus at its original, high-potency baseline value.
- Contextual and Extinction Controls: Sub-groups were monitored to verify that the high density of noise presentations did not induce a state of generalized motor exhaustion, permanent hearing damage, or non-specific contextual conditioning that could artificially depress or distort subsequent task performance.
By the conclusion of Phase Two, the experimental manipulation had produced an organism that possessed a fully intact, historic CS-US associative memory formed in Phase One, but which now evaluated the physical US itself as emotionally neutral and behaviorally irrelevant.
4.3 Phase Three: Critical Extinction Probe Testing
Phase Three was the definitive empirical probe: the unreinforced presentation of the conditioned stimulus to evaluate the status of the conditioned response. The animal was returned to the baseline operant testing chamber and allowed to establish a stable, continuous rate of drinking or lever pressing. Once baseline behavior stabilized, the conditioned stimulus (e.g., the 30-second tone) was presented in complete isolation. No unconditioned stimuli—neither the original loud noise nor any alternate reinforcers—were delivered during this testing phase.
The utilization of an extinction probe was an absolute methodological requirement. Had Rescorla delivered the devalued US at the termination of the CS during testing, any observed change in the animal’s conditioned response could be trivially dismissed by S-R theorists as new, direct learning. In that scenario, the animal would simply be experiencing a new trial of the CS paired with a now-neutral event, triggering rapid first-order counterconditioning or standard extinction. To prove that the behavioral change was driven by an internal cognitive inference, the test had to be conducted without the primary reinforcer present. The animal had to demonstrate its behavioral adjustment on the very first presentation of the CS, drawing exclusively on the internal integration of its prior memories.
During this probe test, Rescorla precisely recorded the behavioral topography of the response. The primary dependent measures included the latency to resume appetitive baseline behavior upon CS onset, the total lick or press rate during the CS interval, and the mathematical suppression ratio. The performance of the devalued group was systematically compared against the non-devalued control group. If the S-R model held true, both groups would exhibit identical, near-zero suppression ratios, indicating robust behavioral suppression driven by the historic habit strength stamped between the CS and the motor output. If the S-S model held true, the devalued group would demonstrate an immediate, significant elevation in their suppression ratios toward $0.50$, indicating that because the internal representation of the US was no longer frightening, the CS had lost its capacity to evoke defensive fear behavior.
5. Techniques for Devaluing the Unconditioned Stimulus
5.1 Acoustic and Defensive Habituation
In Rescorla’s original 1973 experimental designs, habituation to an aversive acoustic unconditioned stimulus served as the primary technique for US devaluation. This methodological selection was strategically brilliant because it bypassed the somatic and systemic physiological complications inherent to nociceptive stimuli. Acoustic habituation relies on the evolutionary capacity of the mammalian nervous system to decrease behavioral responsiveness to repeated sensory inputs that prove to be devoid of biological consequence.
The protocol required exposing the subject to prolonged, repetitive trains of the high-intensity sound (typically a 2-second, 110-dB burst of white noise or a specific horn frequency) delivered over hundreds of trials with short, regularized inter-stimulus intervals. The depth of the resulting habituation was empirically monitored via baseline startle chambers equipped with piezoelectric accelerometers to verify that the motoric flinch and somatic defensive reflexes had diminished to baseline noise levels. This process systematically downscaled the functional activation of defensive brainstem and midbrain circuits without inducing structural tissue injury.
However, the acoustic habituation technique presents specific methodological challenges that require rigorous experimental controls:
- Spontaneous Recovery: Habituation is notoriously susceptible to the spontaneous recovery of the habituated response over time. The temporal window between Phase Two habituation and Phase Three probe testing must be minimized to prevent spontaneous restoration of the US’s unconditioned aversive salience.
- Acoustic Fatigue and Auditory Damage: Researchers must calibrate the decibel levels and duty cycles to ensure that the loss of responsiveness reflects active central nervous system habituation rather than peripheral auditory threshold shifts (temporary or permanent hearing damage).
- Stimulus Generalization: The habituation must be specific to the US itself; if the habituation regimen induces a generalized sensory dampening that blunts the organism’s perception of the conditioned stimulus, the experimental logic collapses.
5.2 Conditioned Taste Aversion (CTA) as Devaluation
While Rescorla pioneered US devaluation using acoustic habituation, the paradigm was substantially expanded and popularized across behavioral neuroscience through the integration of Conditioned Taste Aversion (CTA) protocols. In an appetitive CTA devaluation experiment, the primary unconditioned stimulus is an appetitive nutrient—typically a novel flavored solution, a sucrose pellet, or a maltodextrin liquid. During Phase One, an arbitrary conditioned stimulus (such as a visual light or an auditory tone) is paired with the delivery and consumption of this appetitive US, establishing a robust conditioned approach response (such as rapid magazine entry or food-cup approach behavior).
In Phase Two, the US is independently devalued by pairing its free, unconstrained consumption with the delayed induction of transient gastrointestinal malaise. This visceral illness is chemically induced via an intraperitoneal injection of a mild emetic agent, most commonly lithium chloride (LiCl). The animal consumes the food or liquid outcome in the home cage or a distinct holding environment, receives the LiCl injection, and experiences several hours of nausea. Through internal visceral associative mechanisms, the hedonic status of the US undergoes a profound inversion: a previously craved, highly reinforcing reward is transformed into an intensely aversive substance that elicits robust disgust behaviors (such as gape reactions and chin rubs).
Conditioned Taste Aversion offers immense experimental power because the devaluation effect is enduring, biologically potent, and virtually immune to spontaneous recovery across standard experimental intervals. Furthermore, CTA fundamentally alters the unconditioned stimulus’s hedonic value (its qualitative palatability) rather than merely reducing general drive states. When the CS is subsequently presented in Phase Three, the animal is not merely indifferent to the food; it mentally simulates a substance that is actively repulsive. CTA devaluation provided empirical proof that conditioned stimuli evoke multidimensional sensory representations of outcomes that are linked to hedonic evaluation circuits.
5.3 Specific Satiety Protocols
A third major methodology for US devaluation is sensory-specific satiety. Pioneered by behavioral physiologists and extensively refined within modern cognitive neuroscience, sensory-specific satiety devalues an appetitive unconditioned stimulus by exploiting homeostatic and sensory satiation mechanisms without resorting to toxicological or chemical interventions.
In this protocol, an animal (or human subject) that has learned an association between a CS and a specific food outcome (e.g., Outcome 1: high-sucrose pellets) is placed in an environment immediately prior to the critical test phase and granted unlimited, ad libitum access to that specific food for an extended duration (typically 30 to 60 minutes). The subject consumes the food to total satiety, driving the temporary sensory and hedonic valuation of that specific substance toward zero. Crucially, the subject is not generally sated; if offered an alternative, qualitatively distinct food (e.g., Outcome 2: high-protein pellets or standard chow), the animal will avidly consume it. This isolates the devaluation to the specific sensory identity of the outcome while leaving general metabolic drive relatively unperturbed.
Sensory-specific satiety presents notable advantages and specific experimental constraints:
- Within-Subject Double-Dissociations: Researchers can train two distinct conditioned stimuli paired with two distinct appetitive outcomes (CS1 $\rightarrow$ US1; CS2 $\rightarrow$ US2). By sating the subject exclusively on US1, the experimenter can evaluate the response to CS1 versus CS2 in the identical animal within the same probe session. This provides an internal control against non-specific satiety, motivational lethargy, or motor fatigue.
- Temporal Transience: Unlike the permanent hedonic shift induced by lithium chloride, sensory-specific satiety is fundamentally transient. As digestion proceeds and metabolic clearance occurs, the outcome gradually recovers its hedonic and motivational value over several hours. Experimental probe tests must be executed rapidly within the immediate post-satiation window.
6. Experimental Findings and Quantitative Observations
6.1 Response Attenuation in the Test Phase
The quantitative results of Rescorla’s original experiments—and the hundreds of programmatic replications that followed—provided definitive, mathematically robust proof in favor of S-S representational models. During the critical Phase Three extinction probe trials, subjects in the devalued condition demonstrated a dramatic, statistically significant attenuation of the conditioned response compared to non-devalued controls.
In Rescorla’s acoustic habituation designs, the conditioned suppression ratios revealed an immediate divergence between experimental cohorts. Animals in the control condition, whose loud noise US retained its high aversive valence, showed profound behavioral suppression: upon presentation of the CS, their baseline drinking or lever-pressing rates plummeted, yielding suppression ratios tightly clustered between $0.05$ and $0.15$. In stark contrast, animals whose loud noise US had undergone independent Phase Two habituation demonstrated minimal suppression: upon the onset of the CS, they continued to drink or press with near-normal vigor, registering suppression ratios ranging between $0.38$ and $0.48$.
The temporal dynamics of this response modification were striking: the attenuation of the conditioned response occurred instantaneously, on the very first trial of the test phase. There was no learning curve, no transient period of high suppression followed by rapid within-session extinction, and no transitional phase required for behavioral adaptation. The very first second the CS sounded, the animal executed an operational computation based on the updated value of the US. Because the animal had never experienced the CS in conjunction with the habituated noise, this immediate behavioral emancipation could only occur if the CS evoked an internal representation of the US whose current value was accessed to determine whether defensive suppression was warranted.
6.2 Control Conditions and Verification of Specificity
To establish that the observed response attenuation was driven by the specific representational updating of the US rather than experimental artifacts, Rescorla and subsequent researchers implemented stringent control conditions. The validity of the devaluation paradigm hinges entirely on ruling out alternative non-associative explanations, such as generalized behavioral changes, motor exhaustion, sensory cross-habituation, or contextual conditioning.
One essential control involved the inclusion of non-habituated and pseudo-conditioned baseline cohorts. Researchers compared the devalued group against animals that received identical handling, identical confinement in the habituation apparatus, and equivalent exposure to novel environments without the delivery of the specific US. These controls demonstrated that the elevated response rates in the devalued cohort were not the byproduct of handling stress, generalized arousal shifts, or environmental dishabituation.
A rigorous demonstration of specificity is achieved through double-dissociation designs utilizing multiple stimuli and outcomes within the same organism:
- Experimental Setup: Subject is trained with two distinct stimuli and outcomes:
$$\text{Tone (CS1)} \rightarrow \text{Sucrose Pellets (US1)}$$
$$\text{Light (CS2)} \rightarrow \text{Maltodextrin Liquid (US2)}$$ - Targeted Devaluation: The subject is exposed to an independent devaluation protocol (via CTA or specific satiety) targeting exclusively US1, leaving US2 completely pristine.
- Extinction Probe: In the Phase Three test, both CS1 and CS2 are presented under extinction conditions.
- Empirical Result: The organism exhibits robust, profound attenuation of conditioned approach exclusively to CS1, while maintaining full, unattenuated conditioned responding to CS2.
This within-subject selective devaluation empirically eliminates any generalized explanation based on systemic malaise, sickness-induced motor retardation, depression of general appetite, or non-specific sensory blunting. The behavioral modification is laser-focused on the specific associative channel tied to the devalued outcome representation.
6.3 Residual Responding and Incomplete Suppression
Despite the unmistakable, robust attenuation of conditioned responding characteristic of the devaluation effect, experimental data consistently reveal a nuanced quantitative phenomenon: devaluation rarely results in the total, 100% eradication of the conditioned response. A small but persistent degree of residual responding typically remains. Even in thoroughly devalued animals, the suppression ratio rarely matches the theoretical baseline of unconditioned neutrality ($0.50$), and in appetitive paradigms, animals often emit a low baseline frequency of magazine entries or approach responses to the CS of a devalued reward.
This residual response profile has fueled intense theoretical debate within behavioral psychology. Researchers have advanced three primary interpretations:
- Imperfect US Devaluation: The secondary intervention (habituation or CTA) rarely reduces the subjective biological value of the US to absolute zero. A tiny fraction of residual salience or affective charge remains associated with the US representation, which the CS faithfully extracts and expresses during testing.
- Dual-System Architecture (Concurrent S-S and S-R Encodings): Conditioning does not occur via an exclusive S-S pathway; rather, the nervous system concurrently builds both an S-S representational branch and an S-R direct motor branch. While US devaluation completely eliminates the S-S contribution to the conditioned response, the autonomous S-R habit strength—which was stamped directly into sensorimotor circuits during Phase One—remains immune to the post-training value shift, generating a baseline residue of automated responding.
- The Overtraining Effect: The proportion of residual, devaluation-resistant responding is directly modulated by the amount of initial training. In 1985, researchers demonstrated that if Phase One conditioning is extended over hundreds of additional trials beyond the acquisition asymptote, the conditioned response becomes increasingly resistant to post-conditioning US devaluation. This transition marks the gradual shift from goal-directed representational control to autonomous habitual execution.
7. Cognitive Implications: Mental Representations and Expectancy
7.1 The Stimulus-Stimulus (S-S) Confirmation
The definitive success of Robert Rescorla’s US devaluation experiments delivered the empirical coup de grâce to the assertion that classical conditioning could be explained entirely through peripheral Stimulus-Response mechanisms. By demonstrating that a conditioned response can be substantially modified without altering the physical properties of the CS, without providing additional CS-US training trials, and without directly extinguishing the CS, Rescorla established that learning involves the construction of an internal cognitive architecture.
The operational topology validated by devaluation requires that the CS must evoke an active, editable representation of the US. The behavioral response (CR) is not hard-linked to the sensory input; it is an emergent output calculated at the moment of recall by evaluating the intersection between the activated US representation and the organism’s current motivational state. If the internal representation of the US is charged with high motivational significance, the response executes vigorously; if the representation has been downgraded to affective neutrality or aversion, the response is actively withheld or redirected.
This empirical confirmation redefined the organism within experimental psychology. The animal could no longer be characterized as a passive automaton whose behavior is dictated mechanically by historical patterns of reinforcement. Instead, the organism was shown to be an active, inferential evaluator. Conditioned behavior represents a real-time logical deduction: “Cue A predicts Outcome B; I currently evaluate Outcome B as worthless or repulsive; therefore, I will not waste metabolic energy or endanger myself by executing an approach or defensive response to Cue A.” Classical conditioning was revealed to be a representational, cognitive process operating through associative mechanisms.
7.2 The Nature of Conditioned Representations
The validation of S-S learning immediately raised a deeper, more sophisticated theoretical question: what is the precise qualitative nature of the mental representation evoked by the conditioned stimulus? Research following Rescorla’s breakthrough demonstrated that an unconditioned stimulus representation is not a monolithic, one-dimensional cognitive symbol. Rather, it is a rich, multidimensional neural ensemble comprising distinct sensory, affective, and hedonic features.
Experiments employing nuanced devaluation procedures have demonstrated that these representational dimensions can be functionally dissociated:
- The Sensory-Specific Dimension: This component encodes the objective physical properties of the outcome—its unique taste, texture, spatial location, acoustic frequency, visual pattern, and caloric density. Sensory-specific satiety selectively alters this dimension without damaging general affective processing.
- The Affective-Motivational Dimension: This component encodes the general valence and emotional arousal elicited by the outcome—whether it is broadly good, bad, appetitive, or aversive. Systemic state shifts (such as general stress induction or non-specific drive shifts) can manipulate this dimension independently of sensory identity.
- The Hedonic Palatability Dimension: This component encodes the intrinsic pleasure or disgust elicited by the sensory event (its “liking” rather than its “wanting”). Conditioned Taste Aversion using LiCl selectively inverts this hedonic valuation, converting positive affective reactions into active rejection behaviors.
Furthermore, devaluation research revealed the existence of mediated conditioning and retrospective revaluation. In these paradigms, an unobserved, absent stimulus can be associatively modified because its mental representation has been reactivated by another cue. Representations are not dormant files stored in passive memory registers; they are dynamic, interactively linked nodes that maintain structural coherence across temporal gaps, allowing organisms to continually update their cognitive models of the world as environmental contingencies evolve.
7.3 The Concept of Anticipation and Expectancy
Prior to Rescorla’s work, the terms “anticipation” and “expectancy” were viewed with deep skepticism by the behavioral establishment. They were considered unscientific, anthropomorphic labels that merely redescribed behavioral observations without providing explanatory mechanisms. The US devaluation paradigm changed this status by providing an objective, mathematically rigorous operationalization of cognitive expectancy.
Within the devaluation framework, expectancy is not a vague philosophical state; it is a measurable intervening variable with definitive predictive power. An expectancy exists if and only if an organism’s behavioral response to an antecedent cue varies systematically as a function of the current independent valuation of the anticipated outcome, under conditions where the physical associative bond between the cue and the response has never been directly modified. By satisfying this strict operational criterion, Rescorla formalized expectancy as a concrete functional construct within experimental psychology.
This operationalization cleanly distinguishes between a reflexive reaction and an anticipatory action. A reflexive reaction is backward-looking: it is driven blindly by the physical stimulus energy of the environment interacting with historically established neural pathways. An anticipatory action is forward-looking: it is guided by an internal mental simulation of a prospective future state. US devaluation proved that animals are model-builders. They project themselves forward in time, evaluating current environmental signals through the lens of internal simulations regarding the prospective consequences of those signals, dynamically optimizing their behavior to match the shifting demands of their internal and external worlds.
8. The Interplay Between the Rescorla-Wagner Model and Devaluation
8.1 Algorithmic Bounds of the 1972 Model
A profound historical irony in the history of psychology is that Robert Rescorla co-authored the most famous mathematical model of classical conditioning—the Rescorla-Wagner model (1972)—just one year before publishing the seminal 1973 devaluation experiments that revealed the model’s fundamental algorithmic boundaries. While the Rescorla-Wagner model was revolutionary in its capacity to calculate associative changes based on prediction error, its baseline mathematical architecture cannot natively account for the phenomenon of US devaluation.
The mathematical limitation stems from the model’s core learning equation:
$$\Delta V_{i} = \alpha_{i} \beta (\lambda – \sum V)$$
In this algorithm, the associative strength of a conditioned stimulus ($V_i$) can only change when the stimulus is physically present on a trial, thereby allowing the computation of the prediction error term $(\lambda – \sum V)$. During Phase Two of a US devaluation experiment, the CS is totally absent ($\alpha_{i} = 0$). Consequently, the model dictates that $\Delta V_{i} = 0$. The associative value of the CS remains entirely frozen at its asymptotic level throughout the entire devaluation procedure.
Furthermore, during the critical Phase Three extinction probe test, both the devalued group and the non-devalued control group begin the test trial with identical historical associative values ($V_{CS} = \lambda$). Because no US is delivered during an extinction test ($lambda = 0$), the Rescorla-Wagner model predicts that on the very first test trial, both groups must exhibit an identical, maximal conditioned response, driven by the identical pre-existing value of $V$. The model predicts that any change in responding can only emerge after the organism has experienced non-reinforced trials, driving standard extinction. The empirical reality of US devaluation—where the devalued group exhibits an immediate, profound reduction in responding on Trial 1—exposes the limitation of the standard single-variable delta-rule model, which treats $V$ as a direct associative connection rather than an active retrieval link to a modifiable US node.
8.2 Theoretical Modifications and Hybrid Models
The algorithmic limitations revealed by the devaluation paradigm catalyzed a wave of theoretical innovation, prompting researchers to develop more sophisticated associative models capable of reconciling computational precision with representational flexibility.
One major advancement was Allan Wagner’s Sometimes-Opponent-Processes (SOP) model, and its subsequent affective extension, AESOP. The SOP model replaced the monolithic scalar associative strength $V$ with a dynamic network of representational nodes composed of large sets of constituent elements. Each stimulus—whether a CS or a US—is represented as a node whose elements transition across three distinct computational activation states:
- Primary Active State ($A1$): A high-intensity, short-lived focal activation state triggered by the physical presentation of the stimulus.
- Secondary Active State ($A2$): A lower-intensity, decaying activation state characterized by representational retrieval or rehearsal.
- Inactive State ($I$): The baseline, dormant state of the representational elements.
In the SOP framework, learning occurs when the elements of the CS and US occupy specific overlapping activation states. AESOP expanded this architecture by separating the US node into distinct sensory and affective processing networks. Under this formulation, when an outcome is devalued via habituation or illness, the elements comprising its affective processing network undergo state transitions or acquire inhibitory links. When the CS is presented during testing, it retrieves the US node into the $A2$ state; however, because the affective node’s internal properties have been altered, the behavioral output generated by this representational retrieval is markedly diminished, providing an associative mechanism for the devaluation effect.
Concurrently, hybrid models emerged that integrated the attention-based mechanisms of the Pearce-Hall model with representational networks. In the Pearce-Hall architecture, learning is modulated by the associability of the conditioned stimulus, which is determined by how accurately the outcome was predicted on prior trials. Modern computational adaptations allow the internal representation of the outcome to maintain identity-specific vectors that update independently of the CS’s associative weights, formalizing the distinction between prediction error computation and identity-specific outcome tracking.
8.3 Reconciling Computational Precision with Representational Flexibility
The ultimate theoretical synthesis between Rescorla’s computational modeling and his empirical devaluation findings arrived with the emergence of Reinforcement Learning (RL) frameworks within computational neuroscience, specifically the operational distinction between Model-Free and Model-Based algorithmic architectures.
The classic Rescorla-Wagner model is the mathematical precursor to contemporary Model-Free reinforcement learning (such as the Temporal Difference learning algorithm, $TD(0)$). Model-free algorithms are computationally inexpensive: they learn a single scalar value (a $Q$-value or state value $V$) associated with a state or an action. The algorithm does not encode what specific outcome will occur, what its sensory attributes are, or what causal transitions exist in the environment; it merely tracks the cached historical expectation of reward. Because model-free systems rely on cached values, they are fundamentally devaluation-resistant. If the outcome value changes offline, a model-free agent cannot know this until it physically executes the action, receives the devalued outcome, computes a new negative prediction error, and slowly updates its cached value.
In contrast, Rescorla’s US devaluation experiments empirically validated the biological existence of Model-Based reinforcement learning. A model-based agent constructs an explicit internal cognitive map of the transition probabilities between environmental states and encodes the specific identity and current utility of the terminal outcomes. When presented with a conditioned stimulus, a model-based agent performs an active, prospective tree-search or mental simulation: it retrieves the state transition ($\text{CS} \rightarrow \text{US}$) and multiplies that transition probability by the current, real-time utility assigned to that specific US representation ($U(\text{US})$). When the US is independently devalued in Phase Two, the agent’s internal utility table is updated ($U(\text{US}) \rightarrow 0$). Consequently, when tested with the CS in Phase Three, the model-based agent immediately computes zero expected utility, instantly suppressing the conditioned response without requiring a single prediction error trial.
Today, hierarchical Bayesian frameworks and dual-system computational models treat mammalian behavior as a continuous, dynamic arbitration between model-free (habitual, S-R) and model-based (goal-directed, S-S) computational engines. Robert Rescorla’s devaluation paradigm provided the fundamental behavioral assay used to differentiate, calibrate, and map these dual computational systems in the biological brain.
9. Methodological Variations: Instrumental Versus Classical Devaluation
9.1 Anthony Dickinson’s Instrumental Devaluation Paradigms
While Robert Rescorla designed the US devaluation paradigm to interrogate the associative structure of Pavlovian classical conditioning, British experimental psychologist Anthony Dickinson recognized its revolutionary potential for resolving the nature of instrumental, operant conditioning. Throughout the late 1970s and 1980s, Dickinson and his colleagues at the University of Cambridge adapted Rescorla’s Pavlovian logic into an instrumental framework, creating what is now universally recognized as the primary assay for goal-directed action.
In an instrumental devaluation experiment, the associative relationship under investigation is not between an external CS and a US, but between an organism’s voluntary motor action ($A$)—such as pressing a lever or nose-poking an aperture—and an outcome ($O$), such as a food pellet. Traditional operant theory, rooted in Thorndikian law of effect and Skinnerian radical behaviorism, asserted that instrumental behavior was purely S-R: the delivery of the reinforcer served merely to stamp in a direct connection between the contextual stimuli of the operant chamber and the motor action of pressing the lever.
Dickinson challenged this S-R consensus by applying Rescorla’s three-phase protocol:
- Phase One (Instrumental Acquisition): The animal learns to execute an action to earn a specific outcome ($A1 \rightarrow O1$).
- Phase Two (Outcome Devaluation): The outcome is independently devalued in the absence of the lever, typically via lithium chloride-induced CTA or specific satiety.
- Phase Three (Extinction Probe Test): The animal is returned to the chamber and presented with the lever for an unreinforced probe test.
Dickinson demonstrated that under standard training conditions, animals immediately withheld their lever pressing during the extinction test if that specific action led to the devalued outcome. This established the Action-Outcome (A-O) representational framework. Dickinson formalized the criteria for true goal-directed action: an action is goal-directed if and only if it is mediated by knowledge of the causal contingency between the action and the outcome ($A \rightarrow O$), AND the current subjective value of that specific outcome representation. Dickinson’s instrumental devaluation paradigms directly mirrored Rescorla’s classical findings, confirming that both Pavlovian conditioned reflexes and instrumental voluntary actions are governed by internal outcome representations.
9.2 The Overtraining Effect and Habit Emergence
One of the most consequential discoveries to emerge from the adaptation of Rescorla’s devaluation methodology was the profound impact of extensive training on the associative structure of behavior. While moderately trained instrumental actions are exquisitely sensitive to outcome devaluation (confirming A-O goal-directed control), Dickinson and his colleagues discovered that if animals are subjected to extensive overtraining, the behavior undergoes a fundamental neurocomputational transformation: it becomes entirely resistant to outcome devaluation.
In these classic experiments, rats trained on a moderate schedule of lever pressing (e.g., 100 to 120 total reinforcers) showed immediate suppression of lever pressing following outcome devaluation. However, if the training was extended over hundreds or thousands of additional trials (overtraining), the subsequent devaluation of the outcome via LiCl had zero effect on test performance: the overtrained animals continued to press the lever at maximal, vigorous rates during the extinction test, despite the fact that they completely refused to consume the outcome when presented to them freely in their home cages.
This empirical finding validated the modern dual-process theory of action:
- Goal-Directed Behavior (A-O): Dominates during early learning. It is computationally demanding, relies on explicit internal representations of the outcome, is highly flexible, and is exquisitely sensitive to US/Outcome devaluation.
- Habitual Behavior (S-R): Emerges gradually through repetition and overtraining. It is computationally efficient, automatic, resistant to cognitive distraction, and structurally autonomous from the outcome’s current value. It is fundamentally devaluation-insensitive.
Interestingly, this transition from goal-directed to habitual execution displays marked differences between instrumental and classical paradigms. While instrumental actions readily transition into devaluation-resistant S-R habits following overtraining, pure Pavlovian conditioning displays an extraordinary persistence of S-S representational control. Even after extensive classical overtraining, a conditioned stimulus frequently retains significant sensitivity to US devaluation, reflecting the evolutionary imperative for Pavlovian predictive systems to maintain faithful sensory and hedonic representations of impending biological threats and rewards.
9.3 Pavlovian-to-Instrumental Transfer (PIT) and Devaluation
The convergence of Rescorla’s classical devaluation and Dickinson’s instrumental devaluation reached its methodological pinnacle in the investigation of Pavlovian-to-Instrumental Transfer (PIT). The PIT paradigm interrogates how a purely classical predictive cue (a Pavlovian CS) influences the execution and vigor of an independently trained voluntary instrumental action ($A$).
In a standard PIT experiment, an animal undergoes classical conditioning in one set of sessions ($\text{Tone} \rightarrow \text{Pellet}$) and instrumental conditioning in separate sessions ($\text{Lever Press} \rightarrow \text{Pellet}$). During testing, the lever is made available under extinction conditions, and the tone is unexpectedly presented. The baseline rate of lever pressing reliably surges upon the onset of the tone—the classical cue transfers motivational energy directly into the instrumental motor system. However, integrating US devaluation into this paradigm revealed a crucial bifurcation within PIT itself, dividing the phenomenon into Specific PIT and General PIT:
| Transfer Type | Underlying Associative Mechanism | Sensitivity to US Devaluation | Primary Neural Substrates |
|---|---|---|---|
| Specific PIT | CS retrieves the specific sensory identity of the outcome; biases choice toward the specific action paired with that outcome. | Paradoxically Insensitive: The CS continues to selectively trigger the associated action even after that specific outcome has been devalued via CTA or satiety. | Basolateral Amygdala (BLA), Nucleus Accumbens Shell (NAc Shell). |
| General PIT | CS evokes a general, non-specific affective arousal state; energizes all ongoing instrumental actions indiscriminately. | Highly Sensitive: Devaluing the general outcome or shifting the primary motivational drive state abolishes the general motivational energization. | Central Nucleus of the Amygdala (CeA), Nucleus Accumbens Core (NAc Core). |
The finding that Specific PIT is paradoxically insensitive to outcome devaluation provided profound insights into the architecture of cue-driven action. It revealed that a conditioned stimulus can directly prime the motor system to select a specific action through sensory cue-reactivity, completely bypassing the current hedonic or motivational value of the goal. This behavioral mechanism explains why environmental cues (such as corporate fast-food logos or drug paraphernalia) can automatically trigger specific consumption and seeking actions even in individuals who are completely sated or who have undergone cognitive interventions aimed at devaluing the substance.
10. Neurobiological Substrates of US Representation and Value Updating
10.1 Amygdalar Circuits and Associative Encoding
The behavioral operationalization of the US devaluation paradigm provided neuroscientists with the precise functional assay required to isolate and map the neural circuits responsible for maintaining and updating outcome representations. Over three decades of modern neurobiology have firmly established that the amygdaloid complex is the primary subcortical hub for encoding and updating these representations.
Crucially, neuroscientists have uncovered a profound functional dissociation between the subnuclei of the amygdala:
- The Basolateral Amygdala (BLA): The BLA is the critical neural substrate for encoding the sensory-specific representation of the unconditioned stimulus and linking that representation to predictive cues. Excitotoxic bilateral lesions or optogenetic/pharmacogenetic silencing of the BLA completely abolishes sensitivity to US devaluation. When the BLA is inactivated, an animal undergoing an extinction probe test following devaluation continues to respond to the CS with unattenuated vigor—it behaves like an S-R machine, unable to integrate the post-conditioning value change. Single-unit electrophysiological recordings confirm that individual BLA neurons fire selectively to conditioned cues in a manner that directly mirrors the updated hedonic value of the specific outcome they predict.
- The Central Nucleus of the Amygdala (CeA): In stark contrast to the BLA, the CeA is entirely dispensable for outcome devaluation. Lesions of the CeA leave sensitivity to US devaluation completely intact. Instead, the CeA is dedicated to encoding general affective valence, mediating Pavlovian conditioned orienting responses, and driving General Pavlovian-to-Instrumental Transfer. The BLA encodes what specific outcome is coming, while the CeA encodes whether the impending event is broadly good or bad.
These findings established that the basolateral amygdala is not merely a primitive fear center, as historically characterized, but an advanced computational node responsible for maintaining dynamic, multidimensional representations of biological outcomes and broadcasting those representations to the neocortex for behavioral decision-making.
10.2 Orbitofrontal Cortex (OFC) and Mental Simulation
While the basolateral amygdala encodes the detailed sensory representation of the outcome, the Orbitofrontal Cortex (OFC) serves as the essential neocortical engine that utilizes those representations to compute real-time value expectations and execute cognitive mental simulations. Extensive research by Geoffrey Schoenbaum, Peter Holland, and their contemporaries has demonstrated that the OFC is functionally indispensable for the behavioral expression of the US devaluation effect.
Pre-training or post-training neurotoxic lesions of the OFC—as well as transient pharmacological inactivation via GABA agonists delivered immediately prior to the critical Phase Three test—completely disrupt an organism’s capacity to adjust its behavior following outcome devaluation. When tested with the CS, OFC-lesioned animals exhibit robust conditioned responding to cues predicting devalued outcomes, precisely matching the performance of non-devalued controls. Critically, these OFC-lesioned animals show entirely normal unconditioned responses to the devalued outcome itself: if offered the devalued food directly in their cages, they vigorously reject it. Their sensory processing and their capacity to acquire conditioned taste aversions are fully intact. Their fundamental deficit is purely representational and computational: they cannot retrieve the updated outcome value in absentia to guide behavior when only the predictive CS is present.
The OFC operates via dense, reciprocal neuroanatomical projections with the basolateral amygdala. During an extinction probe test, the presentation of the CS activates sensory neocortical areas, which project to the BLA. The BLA retrieves the specific identity of the associated outcome and transmits this representational signal to the OFC. The OFC integrates this identity signal with the organism’s current homeostatic state, calculating the current subjective economic value of the expected outcome. It then projects downstream to the striatum and motor execution networks to permit or suppress the conditioned response. The OFC provides the neural workspace for the model-based cognitive simulation of unobserved events that Robert Rescorla first deduced via behavioral metrics.
10.3 Striatal Circuitry and Habitual Autonomy
The neural arbitration between devaluation-sensitive goal-directed behavior and devaluation-resistant habitual behavior is physically enacted within the distinct functional subregions of the mammalian striatum. Modern neurobiology has demonstrated that the historical transition from S-S/A-O cognitive representations to S-R automated habits corresponds to a systematic anatomical migration of behavioral control from the medial to the lateral compartments of the dorsal striatum.
The Dorsomedial Striatum (DMS)—homologous to the human caudate nucleus—is the critical striatal hub for goal-directed behavioral regulation. The DMS receives robust glutamatergic projections from the prefrontal cortex, the orbitofrontal cortex, and the basolateral amygdala. Lesions, pharmacological blockades, or optogenetic inhibition of the DMS completely eradicate sensitivity to outcome devaluation, forcing even minimally trained animals into immediate, premature devaluation-resistant habits. The DMS actively encodes the contingency between actions, cues, and their specific outcome representations.
Conversely, the Dorsolateral Striatum (DLS)—homologous to the human putamen—is the primary neural substrate for the emergence, consolidation, and execution of automated, devaluation-resistant S-R habits. The DLS receives dense, direct sensorimotor inputs from primary motor and somatosensory cortices. Inactivation of the DLS in overtrained, habitual animals yields an extraordinary behavioral rescue: it reinstates sensitivity to outcome devaluation. When the DLS is silenced, the overtrained animal ceases to respond like an S-R automaton and reverts instantly to goal-directed, model-based evaluation, demonstrating that the underlying cognitive representation (the A-O / S-S network) is never truly destroyed; it is merely overshadowed and suppressed by the dominant, computationally efficient motor routines running within the DLS circuit.
These striatal pathways are dynamically modulated by dopaminergic signaling originating in the substantia nigra pars compacta ($SNc$) and the ventral tegmental area ($VTA$). Phasic dopamine bursts encode temporal difference prediction errors that facilitate both the construction of model-based transition maps and the progressive stamping-in of model-free sensorimotor synaptic plasticity within the DLS. The balance between DMS goal-directed and DLS habitual processing defines the boundary between cognitive behavioral flexibility and compulsive, autonomous behavioral repetition.
11. Contemporary Applications: Addiction, Obsessive-Compulsive Disorder, and Habits
11.1 Maladaptive Habit Formation in Substance Use Disorders
The theoretical and methodological architecture of US and outcome devaluation has emerged as a premier experimental paradigm for deciphering the pathophysiological mechanisms of psychiatric disorders, particularly substance use disorders (SUD). In contemporary clinical and translational neuroscience, resistance to devaluation is utilized as the primary quantitative behavioral biomarker for the transition from recreational, goal-directed drug-taking to compulsive, automated addiction.
In animal models of addiction, subjects that self-administer drugs of abuse—such as cocaine, methamphetamine, alcohol, or nicotine—exhibit an accelerated, aberrant transition from goal-directed action to devaluation-resistant habits. When a non-drug reward (such as sucrose) is paired with lever pressing, animals typically require weeks of overtraining to develop habitual, devaluation-resistant responding. However, when the self-administered outcome is cocaine or alcohol, the behavior becomes profoundly resistant to devaluation after minimal exposure. Even when the drug outcome is paired with lithium chloride-induced illness, footshock punishment, or direct pharmacological adulteration (such as adding bitter quinine to alcohol), addicted animals continue to vigorously press the lever and seek the substance under extinction conditions.
This persistent devaluation resistance is driven by drug-induced neuroadaptations within frontostriatal and mesolimbic circuits:
- Chronic exposure to elevated dopaminergic surges systematically induces dendritic spine remodeling and synaptic restructuring within the nucleus accumbens and the dorsomedial striatum, impairing the neural machinery required for model-based value updating.
- Concurrently, the drug accelerates neuroplasticity within the dorsolateral striatum, prematurely cementing direct S-R habits that bind drug-associated environmental cues directly to drug-seeking motor sequences.
- Crucially, the subjective conscious valuation of the drug (the cognitive outcome representation) becomes completely decoupled from consumption behavior. The individual may consciously despise the drug, fully recognize its devastating somatic and social consequences, and evaluate its outcome value as negative, yet environmental cues automatically trigger the autonomous S-R motor programs stamped into the DLS.
Modern translational therapies increasingly focus on methods to restore devaluation sensitivity, utilizing targeted cognitive remediation, transcranial magnetic stimulation (TMS) directed at the prefrontal/orbitofrontal cortex, and deep-brain optogenetic stimulation to disrupt habitual DLS circuits and restore goal-directed cognitive control over consumption behaviors.
11.2 Obsessive-Compulsive Disorder (OCD) and Goal-Directed Deficits
Obsessive-Compulsive Disorder (OCD) is characterized by intrusive, distressing thoughts (obsessions) and repetitive, stereotyped mental or physical behaviors (compulsions) that the individual feels driven to perform. While historical psychodynamic and early cognitive formulations viewed compulsions as conscious, intentional actions executed to neutralize the subjective anxiety generated by obsessions, the application of US devaluation paradigms to human clinical cohorts has fundamentally transformed our understanding of OCD etiology.
Pioneering neurocognitive investigations utilizing computerized translational devaluation tasks have revealed that individuals diagnosed with OCD display a profound, systemic impairment in goal-directed behavioral regulation accompanied by a pathological hyperactivity of automated S-R habits. In these paradigms, OCD patients and healthy controls are trained on instrumental tasks where specific actions avoid an aversive outcome (such as a mild electric wrist-shock or a distressing image). Subsequently, the threat is explicitly devalued: the shock electrode is visibly disconnected, the stimulator is turned off, and the participant is explicitly instructed and shown that the aversive outcome can no longer occur.
When tested under extinction conditions, healthy controls immediately cease executing the avoidance action—their behavior is goal-directed and exquisitely sensitive to the threat devaluation. In stark contrast, individuals with OCD continue to repetitively execute the avoidance response, pressing the avoidance key with unattenuated frequency despite consciously knowing and verbalizing that the threat has been neutralized. Functional neuroimaging (fMRI) conducted during these human devaluation protocols reveals severe structural and functional abnormalities within frontostriatal networks: OCD patients demonstrate hypoactivation within the orbitofrontal cortex and the caudate nucleus (DMS equivalent), coupled with an aberrant hyper-connectivity between the premotor cortex and the putamen (DLS equivalent).
These findings demonstrate that compulsions are not merely purposive attempts to reduce anxiety; they are autonomous, hyperactive S-R motor habits that have broken free from cognitive outcome control. The compulsive behavior executes automatically upon the presentation of environmental or internal stimuli, and the patient post-hoc invents an obsessional rationalization to explain their autonomous, habit-driven action. Modern clinical interventions, including Exposure and Response Prevention (ERP), function precisely by forcing the individual to inhabit the stimulus environment without executing the S-R habit, thereby reactivating and retraining the dormant frontostriatal goal-directed circuits responsible for dynamic outcome valuation.
11.3 Metabolic and Eating Disorders
The neurobehavioral principles extracted from Rescorla’s devaluation experiments provide crucial explanatory frameworks for understanding metabolic dysfunction, obesity, and eating disorders such as binge-eating disorder (BED). At the core of healthy energy homeostasis lies the biological mechanism of sensory-specific satiety: the continuous consumption of a specific nutrient drives a rapid, selective devaluation of that food’s sensory representation, terminating consumption while maintaining nutritional openness to diverse dietary sources.
Clinical devaluation research reveals that individuals struggling with obesity and binge-eating pathologies exhibit a marked breakdown in sensory-specific outcome devaluation. In translational laboratory paradigms, participants are exposed to an appetitive conditioning protocol involving distinct food rewards (e.g., chocolate versus savory snacks), followed by a sensory-specific satiety protocol where one food is consumed to fullness. Healthy-weight cohorts show the standard devaluation effect: during subsequent probe tests, conditioned approach responses and subjective cravings for the sated food plummet to near-zero levels. In obese and BED cohorts, however, this devaluation sensitivity is significantly blunted or absent. The subjects continue to demonstrate intense conditioned approach responses, elevated salivation, and high instrumental effort to obtain cues predicting the sated food, despite reporting complete physiological fullness.
This pathology reflects a dissociation between homeostatic metabolic state and conditioned cue reactivity:
- The modern food environment is saturated with hyper-palatable, industrial food cues that act as pervasive Pavlovian conditioned stimuli.
- In vulnerable individuals, these food cues acquire immense conditioned incentive salience that bypasses internal physiological devaluation signals. The sight, smell, or corporate branding of the food activates an autonomous approach response that fails to incorporate the updated, post-ingestive physiological feedback of the body.
- Neuroimaging studies demonstrate that this failure of food devaluation correlates with blunted BLA-OFC functional connectivity and abnormal dopaminergic signaling within the ventral striatum during post-satiety cue exposure.
Addressing metabolic disorders requires the design of cognitive, nutritional, and pharmacological interventions that restore representational flexibility. Therapeutic strategies focus on retraining attentional allocation away from conditioned food cues and utilizing mindfulness-based sensory revaluation protocols that force the explicit cognitive re-evaluation of the internal visceral outcome representation prior to consumption.
12. Lasting Legacy and Modern Critiques of Rescorla’s Devaluation Research
12.1 Epistemological Contributions to Cognitive Psychology
The historical significance of Robert Rescorla’s US devaluation experiments extends far beyond the operational parameters of associative learning theory. Epistemologically, the paradigm established a gold standard for how experimental psychology can rigorously probe, verify, and quantify unobservable internal cognitive states without abandoning the empirical discipline of methodological behaviorism.
Prior to Rescorla, psychology was caught in a false dichotomy: one had to choose between the rigorous, objective, but mechanistically impoverished world of peripheral behaviorism, or the intellectually rich, but scientifically vulnerable world of mentalistic, introspective cognitive speculation. Rescorla destroyed this dichotomy. He proved that internal representations, expectancies, and cognitive evaluations could be isolated with mathematical precision using objective, observable behavioral dependent variables. An animal’s internal mental representation of a loud noise or a food pellet was no longer a philosophical mystery; it was an empirically verifiable computational node whose causal interactions with environmental inputs could be measured, manipulated, and predicted.
Furthermore, Rescorla’s work provided the epistemological foundation upon which modern cognitive neuroscience was constructed. By establishing that the brain builds internal models of the world, encodes predictive relationships through statistical contingency, and dynamically computes behavioral responses based on forward-looking outcome valuations, Rescorla provided the theoretical blueprints that neurophysiologists needed to locate the physical engrams of expectation. His work demonstrated that the associative mechanisms first mapped by Pavlov were not primitive evolutionary relics, but the universal computational language used by mammalian brains to build, navigate, and update their mental models of reality.
12.2 Methodological Limitations and Replication Realities
Despite its towering status, the US devaluation paradigm is characterized by formidable methodological challenges, subtle boundary conditions, and replication variances that have demanded decades of refinement.
A primary experimental challenge centers on the standardization of deep, asymptotic habituation. In acoustic and defensive devaluation paradigms, completely exhausting the biological salience of a potent unconditioned stimulus without inducing generalized sensory fatigue or contextual conditioning requires exceptional experimental control. Minor variations in ambient acoustic resonance, animal strain differences (e.g., between Sprague-Dawley, Wistar, and Long-Evans rats), or baseline anxiety profiles can drastically alter the rate of habituation decay and spontaneous recovery. If an acoustic US spontaneously recovers even a fraction of its original aversive charge during the interval between Phase Two habituation and Phase Three testing, the probe results become muddied by ambiguous residual responding.
Furthermore, marked replication variances exist across different sensory modalities and motivational systems:
- Aversive vs. Appetitive Asymmetries: Classical fear conditioning using electrical footshocks is notoriously resistant to direct post-conditioning habituation devaluation. The biological imperative to avoid tissue-damaging nociceptive pain is so hardwired that animals cannot be safely or reliably habituated to high-intensity shocks. Consequently, fear devaluation typically requires complex counterconditioning protocols or sensory-specific acoustic designs, which present distinct associative interpretations.
- Contextual Interference: If the Phase Two devaluation procedure is executed within the same physical context as the Phase One conditioning or the Phase Three testing, massive contextual conditioning develops. The background cues of the chamber become associatively linked to the altered US, introducing complex configural and inhibitory interactions that obscure whether response changes are driven by direct US devaluation or contextual blocking and overshadowing.
- Spontaneous Recovery and Satiety Decay: In appetitive designs utilizing sensory-specific satiety, the devaluation effect has an exceptionally narrow temporal lifespan. The animal must be tested immediately following consumption, imposing severe temporal constraints on probe trial designs and preventing extended extinction testing across multiple sessions.
Modern behavioral neuroscientists have overcome many of these classical limitations through the integration of optogenetic and chemogenetic technologies. Rather than relying on prolonged acoustic habituation or emetic toxic injections, contemporary researchers can devalue an unconditioned stimulus with millisecond precision by optogenetically silencing specific BLA-to-OFC projections or targeted hypothalamic reward populations at the exact moment of cue presentation. These modern refinements have universally confirmed Rescorla’s foundational qualitative insights while elevating the temporal resolution of representational revaluation to the sub-second scale.
12.3 Future Trajectories in Representational Research
As associative learning theory marches into the mid-twenty-first century, the foundational principles established by Robert Rescorla’s 1973 devaluation paradigm continue to inspire cutting-edge frontiers across multiple scientific domains.
In optical imaging and neural population decoding, neuroscientists utilizing high-density Neuropixels probes and two-photon calcium imaging can now observe the literal physical manifestations of the mental representations Rescorla deduced. Researchers can record from thousands of individual neurons across the basolateral amygdala, the orbitofrontal cortex, and the hippocampus simultaneously while an animal undergoes US devaluation. Scientists can watch in real time as the neural ensemble representing the US undergoes representational drift and hedonic re-encoding during Phase Two, and witness the immediate, predictive reactivation of this modified neural pattern the instant the CS is delivered during the Phase Three probe test. The mental representation is no longer an inferred theoretical construct; it is a visible, dynamic trajectory through biological neural state space.
In the burgeoning field of computational psychiatry, Rescorla’s devaluation methodology is being utilized to map the precise algorithmic parameters that differentiate healthy cognition from psychiatric illness. By fitting hierarchical Bayesian models to human performance data on translational devaluation assays, researchers can extract individualized computational parameters: an individual’s specific learning rate, their model-based versus model-free arbitration weight ($\omega$), their sensory decay rate, and their representational updating speed. These algorithmic metrics provide objective, mechanistic diagnostic profiles that transcend the subjective, descriptive limitations of traditional diagnostic manuals, opening the door to precision psychiatric medicine targeted at specific computational deficits.
Finally, in artificial intelligence and deep reinforcement learning, Rescorla’s insights into model-based representational structures are actively driving the development of next-generation autonomous systems. Contemporary AI agents that combine model-free deep $Q$-networks with model-based tree-search algorithms (such as MuZero and modern world-model architectures) are designed around the fundamental principles validated by the US devaluation paradigm. These artificial networks maintain predictive internal models of their digital environments, evaluate prospective future states against dynamic utility matrices, and demonstrate the same flexible, immediate behavioral recalibration upon outcome revaluation that Robert Rescorla first observed in laboratory rodents over fifty years ago.
Conclusion
The 1973 Unconditioned Stimulus devaluation experiment conducted by Robert Rescorla represents a watershed moment in the intellectual history of behavioral and cognitive science. At a juncture when experimental psychology was constrained by peripheral Stimulus-Response dogmas and mechanistic assertions of drive reduction, Rescorla’s methodological brilliance provided the definitive, unassailable empirical proof that organisms construct rich, dynamic internal representations of the world around them.
By demonstrating that a conditioned response can be instantaneously attenuated through the independent post-training revaluation of an outcome—in the complete absence of the predictive cue—Rescorla dismantled the S-R reflex machine. He proved that Pavlovian conditioning is not an unthinking, mechanical stamping-in of motor habits, but an elegant, representational cognitive system governed by contingency, expectancy, and forward-looking mental simulation. The conditioned stimulus does not act as an autonomous motor switch; it functions as a cognitive key that unlocks an internal representation of the impending event, whose real-time hedonic, sensory, and motivational value determines whether an action will be executed.
The shockwaves of this paradigm shift continue to reverberate across modern science. From the mathematical formalization of model-based reinforcement learning algorithms to the mapping of frontostriatal circuits in the mammalian brain; from deciphering the compulsive neurobiology of addiction and OCD to engineering world models in artificial intelligence—Rescorla’s devaluation paradigm remains the foundational touchstone. Robert Rescorla did not merely invent an experimental technique; he transformed our ontological understanding of the mind, forever establishing that even the simplest behavioral reflex is anchored within an internal universe of cognitive representations.
References
- Dickinson, A. (1985). Actions and habits: The development of behavioural autonomy. Philosophical Transactions of the Royal Society of London. B, Biological Sciences, 308(1135), 67–78. https://doi.org/10.1098/rstb.1985.0010
- Dickinson, A., & Balleine, B. (1994). Motivational control of goal-directed action. Animal Learning & Behavior, 22(1), 1–18. https://doi.org/10.3758/BF03199951
- Garcia, J., Ervin, F. R., & Koelling, R. A. (1966). Learning with prolonged delay of reinforcement. Psychonomic Science, 5(3), 121–122. https://doi.org/10.3758/BF03328311
- Gillan, C. M., Kosinski, M., Whelan, R., Phelps, E. A., & Daw, N. D. (2016). Characterizing a psychiatric symptom dimension related to deficits in goal-directed control. eLife, 5, e11305. https://doi.org/10.7554/eLife.11305
- Gottfried, J. A., O’Doherty, J., & Dolan, R. J. (2003). Encoding predictive odor representations through instruction and associative learning in human orbitofrontal cortex and amygdala. Journal of Neuroscience, 23(17), 6766–6773. https://doi.org/10.1523/JNEUROSCI.23-17-06766.2003
- Holland, P. C., & Rescorla, R. A. (1975). The effect of sample duration and intertrial interval on conditioned inhibition. Journal of Experimental Psychology: Animal Behavior Processes, 1(3), 207–215. https://doi.org/10.1037/0097-7403.1.3.207
- Hull, C. L. (1943). Principles of behavior: An introduction to behavior theory. Appleton-Century-Crofts.
- Ostlund, S. B., & Balleine, B. W. (2007). Orbitofrontal cortex mediates outcome encoding in Pavlovian-to-instrumental transfer. Journal of Neuroscience, 27(18), 4819–4825. https://doi.org/10.1523/JNEUROSCI.5443-06.2007
- Pavlov, I. P. (1927). Conditioned reflexes: An investigation of the physiological activity of the cerebral cortex (G. V. Anrep, Trans.). Oxford University Press.
- Pearce, J. M., & Hall, G. (1980). A model for Pavlovian learning: Variations in the effectiveness of conditioned but not of unconditioned stimuli. Psychological Review, 87(6), 532–552. https://doi.org/10.1037/0033-295X.87.6.532
- Rescorla, R. A. (1968). Probability of shock in the presence and absence of CS in fear conditioning. Journal of Comparative and Physiological Psychology, 66(1), 1–5. https://doi.org/10.1037/h0025984
- Rescorla, R. A. (1973). Effect of US habituation following conditioning. Journal of Comparative and Physiological Psychology, 82(1), 137–143. https://doi.org/10.1037/h0033816
- Rescorla, R. A. (1974). Element-unconditioned stimulus associations in second-order conditioning. Journal of Comparative and Physiological Psychology, 86(5), 901–908. https://doi.org/10.1037/h0036402
- Rescorla, R. A. (1988). Pavlovian conditioning: It’s not what you think it is. American Psychologist, 43(3), 151–160. https://doi.org/10.1037/0003-066X.43.3.151
- Rescorla, R. A., & Wagner, A. R. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In A. H. Black & W. F. Prokasy (Eds.), Classical conditioning II: Current research and theory (pp. 64–99). Appleton-Century-Crofts.
- Schoenbaum, G., Roesch, M. R., Stalnaker, T. A., & Takahashi, Y. K. (2009). A new perspective on the role of the orbitofrontal cortex in adaptive behaviour. Nature Reviews Neuroscience, 10(12), 885–892. https://doi.org/10.1038/nrn2753
- Tolman, E. C. (1948). Cognitive maps in rats and men. Psychological Review, 55(4), 189–208. https://doi.org/10.1037/h0061626
- Wagner, A. R. (1981). SOP: A model of automatic memory processing in animal behavior. In N. E. Spear & R. R. Miller (Eds.), Information processing in animals: Memory mechanisms (pp. 5–47). Lawrence Erlbaum Associates.
- Yin, H. H., & Knowlton, B. J. (2006). The role of the basal ganglia in habit formation. Nature Reviews Neuroscience, 7(6), 464–476. https://doi.org/10.1038/nrn1919