For more than half a century, the study of associative learning was dominated by an intensely mechanical vision of animal and human cognition. Under the classical behaviorist framework inaugurated by Ivan Pavlov and formalized by mid-century Hullian learning theorists, organisms were conceptualized as passive recording devices. In this classic paradigm, when two environmental events occurred within close temporal and spatial contiguity, an associative bond between their internal representations was forged automatically. The strength of this bond was presumed to depend strictly on the number of pairings, the physiological intensity of the stimuli, and the drive state of the organism. The animal did not selectively interrogate its sensory surroundings; rather, the physical energy of the environment stamped itself upon the nervous system through sheer temporal proximity.
This mechanistic consensus suffered its first profound empirical disruption during the late 1960s, most notably with the discovery of the blocking effect by Leon Kamin. By demonstrating that a redundant stimulus fails to elicit conditioning despite dozens of contiguous pairings with a biologically significant outcome, Kamin forced psychology to confront a radical premise: learning is governed by informational value rather than raw temporal contiguity. In the immediate wake of this revelation, Robert Rescorla and Allan Wagner formulated their seminal 1972 model, which preserved the passive, automatic nature of stimulus processing by locating selectivity entirely within the surprisingness of the unconditioned stimulus. Yet, while the Rescorla-Wagner model explained blocking with elegant mathematical economy, it left a gaping theoretical void. It assumed that an organism’s readiness to learn about a conditioned stimulus remained permanently fixed across trials, completely blind to the subject’s prior learning history with that cue.
Enter Nicholas J. Mackintosh. In his landmark 1975 paper published in Psychological Review, titled “A Theory of Attention: Variations in the Associability of Stimuli with Reinforcement,” the British comparative psychologist delivered a decisive intellectual counterweight to the Rescorla-Wagner paradigm. Mackintosh proposed that animals do not possess a uniform, static capacity to learn about the stimuli they encounter. Instead, organisms actively regulate their attention, dynamically modulating the processing bandwidth—termed associability—allocated to specific sensory inputs based on their past predictive track record. Stimuli that prove to be the most reliable, accurate forecasters of biologically significant events gain attentional priority, while redundant, ambiguous, or uninformative cues are systematically filtered out. This comprehensive treatise explores the origins, mathematical mechanics, empirical validations, neurobiological implementations, and enduring computational legacy of Nicholas Mackintosh’s attentional theory of conditioning.
1. Historical Context and Pre-Mackintosh Conditioning Paradigms
1.1 The Behaviorist Foundations of Classical Conditioning
The dawn of modern learning theory emerged from Ivan Pavlov’s pioneering investigations into the digestive reflexes of canines. In the traditional Pavlovian framework, conditioning was conceptualized as a process of physiological reflex substitution. When an arbitrary neutral stimulus, such as a metronome or auditory tone, repeatedly preceded the presentation of an unconditioned stimulus (US) like meat powder, the nervous system established a direct functional connection between the cortical centers processing the conditioned stimulus (CS) and those executing the unconditioned response (UR). The foundational tenet of this early reflexology was temporal contiguity: the closer in time the conditioned stimulus appeared relative to the unconditioned stimulus, the more rapidly and reliably the conditioned reflex developed. There was no theoretical provision for cognitive deliberation, selective attention, or informational evaluation; associative formation was treated as an inevitable biochemical and physiological consequence of overlapping neural excitation.
During the 1930s and 1940s, Clark Hull and Kenneth Spence systematized Pavlovian ideas into rigorous mathematical formulations of neo-behaviorism. Hullian drive-reduction theory posited that learning consisted of the accretion of habit strength, denoted as sHr, which accumulated as a monotonic, negatively accelerating function of reinforced trials. Reinforcement was defined mechanically as the rapid reduction of an organic drive state (such as hunger or thirst). Under this Hull-Spence architecture, every presentation of a stimulus compound alongside drive reduction automatically strengthened the associative connections of all present sensory elements simultaneously. If a light and a tone were sounded together prior to food delivery, both stimuli were assumed to absorb associative habit strength unconditionally, bounded only by their fixed physiological intensities.
The profound limitation of these early contiguity-driven models lay in their utter inability to account for selective cue utilization. In natural environments, organisms are perpetually inundated by an overwhelming cacophony of sensory inputs: the visual patterns of shifting foliage, ambient temperature fluctuations, olfactory scents, and background auditory noise. If temporal contiguity and drive reduction were sufficient to stamp in associations, an organism would rapidly develop debilitating, maladaptive superstitions, binding every co-occurring environmental feature to every subsequent biological reinforcement. Traditional associationism lacked an intrinsic computational mechanism to explain how an animal could disregard a salient sensory event that faithfully accompanied an unconditioned stimulus, yet selectively forge strong associative bonds with a different, seemingly comparable stimulus occurring at the exact same instant.
1.2 The Rise of Cognitive Associative Learning
The conceptual crisis of passive associationism culminated in 1969 with the publication of Leon Kamin’s groundbreaking blocking experiments. Kamin presented laboratory rats with a multi-stage conditioning paradigm that systematically shattered the bedrock assumption that temporal contiguity guarantees associative acquisition. In the initial phase of the experiment, rats were trained to associate an auditory stimulus (a tone, designated as stimulus A) with an aversive footshock until a robust conditioned fear response was established. In the subsequent phase, Kamin presented the subjects with a compound stimulus consisting of the pre-trained tone alongside an entirely novel visual cue (a light, designated as stimulus B). Crucially, the compound AB was paired with the identical footshock. According to traditional contiguity models, stimulus B should have acquired significant conditioned associative strength, as it occurred in precise temporal contiguity with the shock across multiple reinforced trials.
When Kamin subsequently tested stimulus B in isolation, he observed an astonishing result: the rats displayed virtually no conditioned response to the light. The prior training of stimulus A had completely “blocked” conditioning to stimulus B. Kamin reasoned that stimulus B was not conditioned because it provided no new information; the shock was already fully predicted by stimulus A. To explain this phenomenon, Kamin introduced the informational hypothesis, asserting that an unconditioned stimulus must be surprising if it is to instigate associative learning. If an outcome is already anticipated based on existing predictors, the animal experiences no cognitive disruption, no unconditioned stimulus processing takes place, and redundant concurrent cues are disregarded.
Kamin’s discovery elevated the concept of stimulus competition to the forefront of animal cognition. It became undeniable that sensory elements presented within a compound do not learn independently; rather, they engage in a competitive struggle for associative capture. This empirical breakthrough precipitated a paradigm shift away from peripheral, reflex-based behaviorism toward a cognitive view of associative learning. Organisms were increasingly recognized as active information-processing agents that constantly construct internal predictive models of their environment, prioritizing cues that carry high informational utility while disregarding those that merely offer redundant temporal correlation.
1.3 The Precursor: Rescorla-Wagner Model and Its Constraints
To provide a formal, quantifiable mechanism for Kamin’s findings, Robert Rescorla and Allan Wagner formulated what would become the most influential model in the history of associative learning: the Rescorla-Wagner model of 1972. The genius of the model lay in its mathematical formulation of error-correction, where the change in associative strength of a given stimulus ($\Delta V_i$) on any trial was driven by the discrepancy between the actual biological outcome received ($lambda$) and the aggregate associative prediction of all cues present on that trial ($\Sigma V$). The formal equation was expressed as:
$$\Delta V_i = \alpha_i \beta (\lambda – \Sigma V)$$
In this framework, $\alpha_i$ represents the fixed physical salience of conditioned stimulus $i$, $\beta$ represents the learning rate parameter determined by the unconditioned stimulus, and $(\lambda – \Sigma V)$ represents the prediction error, or US surprise. The model accounted for blocking effortlessly: during the compound AB trials, because stimulus A had already been trained to the reinforcement asymptote ($lambda$), the compound prediction $\Sigma V = V_A + V_B$ was equal to $lambda$. Consequently, the prediction error was zero ($(\lambda – \Sigma V) = 0$), preventing any increment in the associative value of stimulus B ($\Delta V_B = 0$).
Despite its unprecedented predictive triumphs across phenomena like overshadowing, conditioned inhibition, and extinction, the Rescorla-Wagner model possessed a glaring, fundamental architectural constraint: it assumed that the stimulus associability parameter, $\alpha$, was an invariant, non-plastic property of the physical stimulus itself. A tone of 80 decibels had an immutable $\alpha$ value dictated exclusively by the sensory thresholds of the animal’s peripheral nervous system. The model insisted that learning was modulated exclusively from the outcome side—governed by whether the US was surprising—while stimulus input processing remained permanently uniform.
This static treatment of stimulus associability rendered the Rescorla-Wagner model utterly incapable of explaining crucial empirical phenomena, most notably latent inhibition and learned irrelevance. If an animal is repeatedly exposed to a neutral stimulus alone without reinforcement prior to conditioning, its subsequent rate of conditioning to that stimulus is profoundly retarded. Because no US is present during the pre-exposure phase, $lambda = 0$ and $\Sigma V = 0$, producing a prediction error of zero. Thus, according to the Rescorla-Wagner equation, no change in associative strength can occur, and $\alpha$ remains untouched. The model predicted that pre-exposed stimuli should condition at the identical rate as completely novel stimuli—a prediction that was repeatedly, dramatically refuted by empirical data. It was precisely this failure to accommodate dynamic, experience-dependent alterations in stimulus processing that motivated Nicholas Mackintosh to formulate his attentional theory.
2. Theoretical Framework of Mackintosh’s 1975 Attentional Model
2.1 Core Tenets of the Attentional Hypothesis
Nicholas Mackintosh approached associative learning from an epistemological perspective profoundly influenced by comparative ethology and cognitive psychology. In his 1975 treatise, Mackintosh posited that the cognitive bottleneck in conditioning does not lie solely within the processing of the unconditioned stimulus, but fundamentally within the processing of the conditioned stimulus. He asserted that the associability of a conditioned stimulus—denoted by the parameter $\alpha$—is not a static physical constant, but a dynamic, experience-dependent psychological variable. Organisms, Mackintosh argued, possess a finite capacity for processing sensory information; they cannot attend equally to every environmental input.
The core postulate of Mackintosh’s attentional hypothesis is that animals actively allocate attention to stimuli that are reliable, valid predictors of biologically meaningful events. If a stimulus consistently signals an outcome with greater precision than any other concurrent sensory cue, the animal increases the attentional processing power devoted to that stimulus. Conversely, if an environmental cue is redundant, ambiguous, or a poorer predictor of reinforcement relative to other available cues, the organism actively downregulates its attention toward that cue, causing its associability to decline.
Crucially, this attentional reallocation is not simply a passive consequence of associative strength accumulation; it is an active mechanism of selective cognitive gating. The Mackintosh model asserts that learning is an active exploratory process wherein animals form hypotheses regarding which features of their sensory environment possess predictive validity. Rather than being mere associative reservoirs, conditioned stimuli are actively interrogated for their informational efficiency. The central innovation of Mackintosh’s model was to make the rate of future learning about a cue a direct function of its past predictive success, establishing a bidirectional feedback loop between associative knowledge and perceptual processing.
2.2 The Concept of Stimulus Salience versus Associability
To construct a rigorous attentional theory, Mackintosh drew a vital conceptual distinction between two parameters that had previously been conflated in the behaviorist literature: physical stimulus salience and acquired stimulus associability. In historical formulations, salience referred to the intrinsic, unconditioned sensory loudness of an environmental event—the physical decibel level of an auditory tone, the lux rating of a visual flash, or the concentration of a chemical gustatory solution. This physical salience unquestionably imposes boundary constraints on sensory transduction within the peripheral nervous system, determining whether an animal can detect a cue above background environmental noise.
Mackintosh introduced associability as a fundamentally separate, central psychological construct. While physical salience is fixed by hardware constraints, acquired associability represents software-level attention. It reflects the degree of central processing, cognitive representation, and neural bandwidth granted to a specific sensory channel based on associative experience. Two stimuli of identical physical salience can possess profoundly divergent associabilities: a quiet, barely audible clicker can command massive associability if it has historically served as an infallible predictor of predatory danger or food availability, while a blinding, physically intense light can suffer near-total attentional suppression if it has proven to be an irrelevant, uninformative sensory artifact.
From an evolutionary perspective, the adaptive value of this selective attentional filtering is immense. In natural habitats, an animal is perpetually immersed in high-dimensional, noisy sensory arrays. If an organism dedicated equal computational bandwidth to every high-salience sensory perturbation, its cognitive architecture would quickly become overwhelmed by sensory distraction. By selectively elevating the associability of valid environmental forecasters and aggressively dampening the associability of redundant or useless stimuli, the Mackintoshian organism optimizes its metabolic and computational resources, ensuring that synaptic plasticity is reserved exclusively for environmental cues that matter for survival.
2.3 The Rule for Updating Associability
The mathematical and conceptual heart of Mackintosh’s 1975 theory lies in its formal operational rule governing the modification of $\alpha$. Unlike models that rely solely on whether an outcome was expected or unexpected, Mackintosh made the trial-by-trial update of a cue’s associability conditional upon its relative predictive validity. An organism evaluates all stimuli present on a given trial, comparing the predictive accuracy of each individual cue against the collective predictive accuracy of all other concurrent cues.
Conceptually, Mackintosh stated that the change in associability for a given stimulus, $\Delta \alpha_A$, is positive if stimulus A predicts the unconditioned stimulus more accurately than any other stimulus present on that trial. Conversely, the change in associability, $\Delta \alpha_A$, is negative if stimulus A predicts the outcome less accurately than, or no better than, the remaining stimuli in the sensory array. The operational logic can be conceptualized as an ongoing comparative ranking:
- Increment Condition: $\Delta \alpha_A > 0$ if $| lambda – V_A | < | lambda - V_X |$
- Decrement Condition: $Delta alpha_A < 0$ if $| \lambda – V_A | \geq | \lambda – V_X |$
Where $V_A$ represents the associative strength of stimulus A, $V_X$ represents the associative strength of all other concurrent stimuli, and $lambda$ represents the asymptote of reinforcement. This elegant principle dictates that attention does not simply rise or fall in a vacuum; it is governed by comparative predictive power. An organism does not merely ask, “Does this cue predict an outcome?” It asks, “Does this cue predict the outcome better than the other cues available to me right now?” If a cue is superior, attention concentrates upon it. If a cue is inferior or redundant, attention is withdrawn, insulating the animal against unnecessary associative processing.
3. Mathematical Formulation and Mechanics of Associability
3.1 The Mathematical Architecture of the Mackintosh Model
To fully formalize the mechanics of the attentional model, Mackintosh established a dual-equation system. The first equation defines the trial-by-trial change in associative strength ($\Delta V$) for a specific stimulus, preserving the classic linear error-correction format popularized by Rescorla and Wagner, but introducing the critical modification that $\alpha$ is dynamic and stimulus-specific:
$$\Delta V_A = \alpha_A \beta (\lambda – V_A)$$
Notice a crucial distinction in the error-correction term: in Mackintosh’s original 1975 formulation, the prediction error for stimulus A is frequently calculated relative to the individual associative strength of stimulus A ($(\lambda – V_A)$), rather than the strictly compounded sum of all cues ($\Sigma V$), though compound variants were explored to account for complete asymptotic overshadowing. Here, $\beta$ remains the rate parameter associated with the unconditioned stimulus, and $lambda$ represents the maximum associative capability supported by the US magnitude.
The revolutionary component of the model resides in the second equation: the associability update algorithm. The parameter $\alpha_A$ updates dynamically according to the following behavioral conditions:
$$\Delta \alpha_A > 0 \quad \text{if} \quad |\lambda – V_A| < |\lambda - V_X|$$
$$\Delta \alpha_A < 0 \quad \text{if} \quad |\lambda - V_A| \geq |\lambda - V_X|$$
Here, $V_X$ denotes the associative strength of all other cues present on the trial except stimulus A (or, in multi-cue implementations, the predictive validity of the best competing alternative stimulus). If the absolute error between the reinforcer and stimulus A is lower than the error between the reinforcer and competing cues, stimulus A was the superior predictor, and $\alpha_A$ increases by a step parameter (often denoted as $c(1 – \alpha_A)$ to impose an upper asymptote at 1.0). Conversely, if the predictive error of stimulus A is equal to or greater than the predictive error of competing cues, $\alpha_A$ decreases by a decay parameter (e.g., $c(\alpha_A – 0)$ to impose a lower bound at 0.0).
3.2 Quantitative Implementation of Relative Predictiveness
The mathematical operation of relative predictiveness establishes an algorithmic thresholding dynamic that fundamentally shapes trial-by-trial learning trajectories. Rather than treating sensory cues as passive elements that continually absorb associative weight until a global outcome error is eliminated, Mackintosh’s system functions as an informational sieve. The model continually audits the performance of each sensory channel against its competitors.
Consider an experiment where a compound of two stimuli, Light (L) and Tone (T), is repeatedly paired with a shock. If the physical conditions or prior training make Light slightly more predictive of the shock than Tone, such that $|lambda – V_L| < |lambda - V_T|$, the mathematical inequality is satisfied in favor of the Light. On t\hat trial,$alpha_L$ increases, while $\alpha_T$ decreases. On the very next trial, because $\alpha_L$ has expanded, Light absorbs an even larger share of associative strength ($\Delta V_L$), further reducing its individual prediction error $|\lambda – V_L|$. Simultaneously, because $\alpha_T$ has degraded, Tone’s capacity to acquire associative strength is suppressed.
This creates a positive feedback loop: the superior predictor becomes progressively more associable, dominating cognitive processing, while the inferior predictor undergoes rapid attentional extinction, its $\alpha$ plunging toward zero. In stable, steady-state environments, this ensures rapid convergence upon the single most reliable predictor. In dynamic environments characterized by probabilistic contingency shifts, the model exhibits complex non-linear behaviors: if contingencies reverse and the formerly reliable cue begins generating massive prediction errors, competing cues whose errors are now comparatively smaller can experience a resurgence in associability, re-opening the attentional gate.
3.3 Computational Simulation of Trial-by-Trial Dynamics
To observe the Mackintosh model in action, computational simulations track the trajectories of associative strength ($V$) and associability ($\alpha$) simultaneously across sequential trials. In a classic compound conditioning simulation involving an overshadowing design—where a high-salience stimulus (A) and a low-salience stimulus (B) are presented simultaneously with a reinforcer—the simulation reveals an immediate divergence in attentional parameters.
On trial 1, both stimuli begin with baseline associabilities, but physical salience gives cue A a modest initial advantage in associative acquisition ($\Delta V_A > \Delta V_B$). By trial 2, cue A’s error $|\lambda – V_A|$ is already strictly smaller than cue B’s error $|\lambda – V_B|$. The Mackintosh update rule triggers: $\alpha_A$ rises toward its ceiling, while $\alpha_B$ begins a steep downward trajectory. Over successive compound presentations, despite the fact that stimulus B is consistently present when reinforcement occurs, stimulus B’s $\alpha$ asymptotically approaches zero. Long before $V_B$ can accumulate meaningful associative strength, the complete collapse of $\alpha_B$ renders stimulus B functionally invisible to the learning machinery.
This computational feature grants the Mackintosh architecture extraordinary resilience against catastrophic interference. In artificial neural networks and connectionist models lacking selective attentional gating, the continuous introduction of redundant or irrelevant sensory features often corrupts established representations—a phenomenon known as catastrophic forgetting or interference. By incorporating dynamic, cue-specific $\alpha$ vectors that selectively throttle processing to uninformative inputs, the Mackintosh model demonstrates how biological systems protect critical predictive representations from degradation by sensory clutter.
4. The Definitive Empirical Experiments by Nicholas Mackintosh
4.1 Mackintosh’s Seminal 1975 Experimental Design
To substantiate his theoretical claims, Nicholas Mackintosh executed a series of brilliant experiments designed to empirically demonstrate that an organism’s rate of learning about a stimulus changes as a direct consequence of that stimulus’s prior predictive validity. Mackintosh realized that to decisively prove his theory over the Rescorla-Wagner model, he had to engineer a paradigm that could isolate variations in stimulus associability ($\alpha$) from differences in net associative strength ($V$). This required an elegant multi-stage discriminative learning design using avian and rodent subjects.
In a prototypical experiment, subjects (such as laboratory rats or pigeons) were trained in a three-stage conditioning sequence. In Stage 1, animals were exposed to compound visual stimuli consisting of two distinct dimensions—for instance, color (stimulus A: red, stimulus B: green) and line orientation (stimulus C: horizontal lines, stimulus D: vertical lines). The training was structured as a conditional discrimination: stimuli were presented as compounds (e.g., Red-Horizontal, Red-Vertical, Green-Horizontal, Green-Vertical). Crucially, the reinforcement contingencies were arranged such that one dimension was perfectly predictive of food delivery, while the other dimension was completely non-predictive.
In the “Predictive” condition, color was the reliable predictor: whenever Red was present (whether paired with Horizontal or Vertical lines), food was delivered; whenever Green was present, no food was delivered. Line orientation was completely uncorrelated with reinforcement. In Stage 2, Mackintosh introduced a new learning problem: both the original predictive cue (color) and the irrelevant cue (orientation) were paired with a novel outcome or placed in a novel discrimination where the previously irrelevant cue was now the target of learning. Stage 3 measured the rate of new acquisition. Mackintosh demonstrated that animals acquired associations to a previously predictive dimension significantly faster than to a previously non-predictive dimension, even when the associative strength of the cues was meticulously equated or driven to zero prior to the test. This confirmed that the predictive history of a stimulus permanently altered its intrinsic associability.
4.2 The Overshadowing and Relative Validity Experiments
The conceptual bedrock of Mackintosh’s experimental program was deeply informed by the relative validity experiments pioneered by Allan Wagner and colleagues in 1968, which Mackintosh extensively refined and interpreted through his attentional framework. In these paradigms, the critical question was whether a conditioned stimulus’s ability to gain associative control was dictated by its objective, absolute correlation with the US, or by its relative correlation compared to alternative concurrent stimuli.
Wagner presented two groups of animals with a target stimulus, a visual cue denoted as $X$. In the “True Relative Validity” group, cue $X$ was presented in compound with either stimulus $A$ or stimulus $B$. Whenever the compound $AX$ appeared, reinforcement occurred; whenever $BX$ appeared, reinforcement was withheld. Thus, stimulus $A$ was 100% predictive of the US, stimulus $B$ was 0% predictive, and stimulus $X$ was reinforced on 50% of its total occurrences. In the “Uncorrelated” control group, compounds $AX$ and $BX$ were both reinforced on exactly 50% of their presentations. Crucially, in both groups, target cue $X$ experienced the exact same absolute schedule of reinforcement: 50% contiguous pairings with the US.
Traditional contiguity theory predicted that cue $X$ should acquire identical associative strength in both conditions. The empirical results, however, were decisive: subjects in the True Relative Validity group developed virtually no conditioned responding to stimulus $X$, whereas subjects in the Uncorrelated group acquired substantial conditioned responding to cue $X$. Mackintosh explained this through dynamic associability: in the True Relative Validity group, stimulus $A$ and stimulus $B$ were perfect, deterministic predictors ($|\lambda – V_A| = 0$ and $|0 – V_B| = 0$), massively outperforming stimulus $X$. As a direct mathematical consequence of Mackintosh’s update rule, $\alpha_X$ was systematically decremented to zero. Stimulus $X$ was overshadowed not because it lacked correlation with the US, but because competing cues possessed superior relative predictive validity, causing the animal to actively withdraw attention from the inferior predictor.
4.3 Measurement Techniques for Attentional Allocation
A central methodological hurdle for Mackintosh was developing operational measures that could directly index covert selective attention without being confounded by overt motor performance. In animal behavior, one cannot simply ask the subject where its attention is focused; attention must be inferred through rigorous, empirically verifiable behavioral proxies.
One primary measurement technique was the evaluation of the rate of subsequent re-acquisition. If an experimental manipulation selectively reduces the associability ($\alpha$) of a stimulus, that stimulus should exhibit significant behavioral inertia when placed in a new learning contingency: it should take many more trials for the animal to acquire conditioned responding to that cue compared to a novel stimulus of identical physical salience. By assessing the slope of the subsequent acquisition curve, Mackintosh successfully separated associative strength (which determines immediate performance on trial 1) from associability (which dictates the velocity of learning across trials 1 through $N$).
A second, highly complementary methodology utilized behavioral orienting responses. In avian and canine conditioning, subjects frequently exhibit overt physiological orientations toward informative stimuli—such as visual fixation, head turns, or direct pecking directed at the localized cue source (sign-tracking). Mackintosh and his contemporaries showed that orienting responses rapidly align with the most predictive elements of a compound stimulus. Furthermore, to definitively prove that variations in learning rates were driven by cognitive attentional mechanisms rather than peripheral receptor fatigue or simple sensory habituation, Mackintosh integrated sophisticated control conditions. He demonstrated that decrements in associability were strictly cue-specific, transferred across diverse unconditioned stimuli, and could be instantly reversed by introducing novel predictive contingencies—outcomes fundamentally incompatible with biological fatigue.
5. Blocking and Unblocking Phenomena Under the Attentional Lens
5.1 Re-examining Kamin’s Blocking Paradigm
The discovery of Kamin’s blocking effect was the catalyst that ignited the cognitive revolution in animal conditioning, and it became the supreme battleground between the Rescorla-Wagner model and Mackintosh’s attentional framework. The two theories approached the same empirical reality from fundamentally opposing theoretical architectures. Under the Rescorla-Wagner model, the failure of the redundant stimulus $B$ to condition during compound $AB$ trials is caused entirely by an absence of unconditioned stimulus surprise: because stimulus $A$ already fully predicts the shock ($\Sigma V = \lambda$), the outcome is expected, the US processing machinery disengages, and no prediction error is generated to support learning.
Mackintosh offered a radically different, input-driven interpretation. He argued that during the compound $AB$ trials, the animal detects both stimulus $A$ and stimulus $B$. However, because the animal has already established a robust associative representation for stimulus $A$, it evaluates the relative predictive validity of both cues on the very first compound trial. Stimulus $A$ possesses an associative strength ($V_A$) close to $lambda$, meaning its predictive error $|\lambda – V_A|$ is near zero. Stimulus $B$, being entirely novel, possesses an associative strength ($V_B$) of zero, yielding a massive predictive error $|\lambda – V_B| = \lambda$.
According to Mackintosh’s associability update rule:
$$|\lambda – V_B| > |\lambda – V_A| implies \Delta \alpha_B < 0$$
Stimulus $B$ is identified as an inferior predictor compared to the established validity of stimulus $A$. As a consequence, the animal rapidly and actively downregulates $\alpha_B$. The blocking of stimulus $B$ does not occur because the US is magically rendered incapable of reinforcing associations; rather, it occurs because stimulus $B$ suffers an catastrophic collapse in its associability. The animal ceases to attend to stimulus $B$, gating it out of central processing. Mackintosh’s blocking is fundamentally an attentional phenomenon: the blocked cue is ignored because it is redundant.
5.2 Unblocking Through Novelty and Contingency Shift
To differentiate between the Rescorla-Wagner US-surprise account and his own CS-associability model, Mackintosh and his contemporaries leaned heavily on the phenomenon of unblocking. If an organism has been trained that stimulus $A$ predicts a specific unconditioned stimulus, and is then presented with compound $AB$, standard blocking ensures that stimulus $B$ fails to condition. However, if the experimenter alters the unconditioned stimulus during the compound phase, blocking is disrupted—a phenomenon known as unblocking.
Unblocking can be triggered through quantitative changes in reinforcer magnitude. If stimulus $A$ was initially trained with a single 1.0 mA footshock, and the compound $AB$ phase subsequently introduces a double footshock (2.0 mA), the unconditioned stimulus is surprising. Rescorla and Wagner easily explained this: the asymptote $lambda$ increases, creating a positive US prediction error $(\lambda_{new} – V_A) > 0$, allowing stimulus $B$ to absorb associative strength. However, Mackintosh pointed to a far more profound empirical test: qualitative unblocking.
In qualitative unblocking experiments, the physical magnitude of reinforcement remains strictly identical, but the qualitative nature of the US is changed. For instance, in an appetitive conditioning setup, stimulus $A$ might predict a sucrose pellet, while the compound $AB$ predicts a polycose pellet of equivalent caloric and reward value. Because the absolute magnitude of reinforcement is unchanged, a pure non-attentional US error-correction model predicts that no surplus associative capacity is available for stimulus $B$. Yet, animals robustly learn about stimulus $B$ under qualitative shifts. Mackintosh demonstrated that the qualitative discrepancy introduces an informational mismatch: stimulus $A$ is no longer an accurate predictor of the specific qualitative features of the outcome. Consequently, the relative predictive validity of stimulus $A$ drops, the downward pressure on $\alpha_B$ is immediately alleviated, and attention is actively re-engaged, permitting rapid associative acquisition to the novel cue.
5.3 One-Trial Blocking and Attentional Dynamics
A crucial line of empirical contention revolved around the temporal dynamics of the blocking effect, specifically the phenomenon of “one-trial blocking.” The Rescorla-Wagner model fundamentally predicts that blocking can never be complete on trial 1 of compound conditioning if stimulus $A$ is trained to anything less than 100% of the asymptote $lambda$. If $V_A$ has reached 0.90 of a 1.0 asymptote, there remains an unassigned 0.10 error window. On trial 1 of compound conditioning, stimulus $B$ must mathematically absorb its proportional share of that remaining 0.10 associative strength.
Mackintosh challenged this premise by analyzing the earliest stages of compound exposure. If selective attention can be modulated rapidly, then a single compound trial might be sufficient for the animal to detect that stimulus $B$ provides zero informational increment over stimulus $A$. Empirical investigations in fear conditioning paradigms revealed that under high-salience conditions, significant blocking effects emerge after a single compound presentation.
This rapid attentional filtering demonstrated that the downward revision of $\alpha_B$ can occur almost instantaneously. The animal does not require dozens of trials of sluggish error accumulation to disregard a redundant cue. Instead, the comparative processing of the sensory array allows the nervous system to make swift, discrete determinations regarding informational relevance. This dynamic theoretical reconciliation established that while associative accumulation ($V$) requires sustained reinforcement across trials, attentional gating ($\alpha$) can shift rapidly, protecting the cognitive system from expending resources on redundant environmental elements.
6. Latent Inhibition and Learned Irrelevance Experiments
6.1 The Mechanism of Latent Inhibition
The definitive empirical superiority of Mackintosh’s attentional framework over purely outcome-driven error-correction models is nowhere more glaring than in the explanation of latent inhibition. First documented by Robert Lubow and A.U. Moore in 1959, latent inhibition refers to the robust, universal finding that repeatedly presenting a neutral stimulus in isolation, without any reinforcement, significantly retards the subject’s subsequent ability to establish a conditioned response to that stimulus when it is later paired with an unconditioned stimulus.
As previously highlighted, the Rescorla-Wagner model suffers a catastrophic mathematical failure when confronted with latent inhibition. During non-reinforced pre-exposure of stimulus $A$, no unconditioned stimulus is present ($lambda = 0$). The animal has no prior expectation of reinforcement ($V_A = 0$). Thus, the prediction error is exactly zero:
$$\Delta V_A = \alpha_A \beta (0 – 0) = 0$$
Because the prediction error is zero, the associative strength remains permanently locked at zero, and because $\alpha_A$ is an invariant constant in the Rescorla-Wagner model, the stimulus remains completely pristine. When reinforcement commences in Phase 2, the stimulus should condition just as rapidly as a completely novel cue. This prediction is unequivocally false; across hundreds of experiments in virtually every vertebrate taxon, pre-exposed stimuli exhibit severe conditioning deficits.
Mackintosh’s model explains latent inhibition with effortless elegance. During the pre-exposure phase, the animal experiences stimulus $A$ within an ambient experimental context (denoted as $C$). The stimulus predicts no biologically salient outcome; in fact, the context predicts the absence of reinforcement just as well as stimulus $A$ does. More fundamentally, because stimulus $A$ is repeatedly followed by nothing, its predictive validity regarding any meaningful environmental change is demonstrably zero. Under Mackintosh’s update rule, when a stimulus has no predictive utility, its associability is systematically degraded:
$$\Delta \alpha_A < 0$$
By the time the pre-exposure phase concludes, $\alpha_A$ has plummeted from its default baseline toward zero. When Phase 2 conditioning begins, the formula governing associative accumulation is $\Delta V_A = \alpha_A \beta (\lambda – V_A)$. Even though the prediction error $(\lambda – V_A)$ is now massive, the multiplier $\alpha_A$ is profoundly suppressed. Consequently, the increments in associative strength ($\Delta V_A$) are tiny on each trial, producing the sluggish, severely retarded acquisition curve characteristic of latent inhibition.
6.2 The Architecture of Learned Irrelevance
While latent inhibition demonstrates that non-reinforced exposure dulls stimulus processing, Nicholas Mackintosh and his students expanded this concept into an even more formidable empirical paradigm: the architecture of learned irrelevance. First systematically investigated by Mackintosh in 1973, learned irrelevance involves exposing animals to a pre-exposure schedule in which both the conditioned stimulus and the unconditioned stimulus are presented, but with an explicitly zero contingency—they occur entirely at random with respect to one another.
In a typical learned irrelevance protocol, an animal receives randomized presentations of a tone and randomized deliveries of food pellets or shocks, such that the occurrence of the tone provides zero predictive information about whether the US will appear, be withheld, or change in frequency. Mackintosh compared the effects of this uncorrelated pre-exposure to both latent inhibition (CS alone) and context-only exposure. The empirical findings were stark: zero-contingency pre-exposure generated an associative retardation deficit that was vastly deeper, more profound, and more resistant to extinction than simple latent inhibition.
Mackintosh’s theoretical explanation was precise: in latent inhibition, the animal learns that the CS predicts nothing. In learned irrelevance, the animal explicitly learns that the CS is irrelevant to the occurrence of the US. Throughout the random presentations, the animal continuously attempts to correlate environmental inputs with the arriving outcomes. On every trial where the US arrives, competing contextual cues predict the US just as well as (or better than) the completely uncorrelated target CS. The comparative predictive validity calculation repeatedly penalizes the target cue. The nervous system computes an active, learned withdrawal of attention, driving the associability parameter $\alpha$ to the absolute floor of cognitive processing. Learned irrelevance represents a profound demonstration of an animal learning an abstract statistical contingency: the formal independence of two events.
6.3 Contextual Modulation of Pre-exposure Deficits
A critical dimension of the empirical literature surrounding latent inhibition and learned irrelevance concerns the role of physical context. If an animal is pre-exposed to a tone in experimental Context 1 (e.g., a dark, steel-grated chamber smelling of vinegar), and subsequent CS-US conditioning is conducted in the identical Context 1, latent inhibition is expressed at its maximum magnitude. However, if the animal is transferred to Context 2 (e.g., a brightly lit, plastic-walled chamber smelling of peppermint) for the conditioning phase, the latent inhibition effect is markedly attenuated—the animal conditions to the tone almost as if it were a novel stimulus.
This contextual specificity ignited a fierce theoretical debate regarding the underlying mechanism of associability decay. Proponents of retrieval-failure models (such as Ralph Miller) argued that latent inhibition does not represent a true loss of stimulus associability; rather, the animal forms an inhibitory CS-nothing association during pre-exposure, which is tied to the pre-exposure context and subsequently interferes with the retrieval of the excitatory CS-US association during testing. Under this view, changing the context merely disrupts the retrieval of the competing memory, restoring performance without any need for attentional constructs.
Mackintosh and his intellectual heirs counter-argued that context acts as an essential conditional gatekeeper for attentional tuning. The associability of a stimulus is not an isolated, context-free property; an animal learns that a cue is irrelevant within a specific ecological setting. When plunged into a completely novel environment, the predictive assumptions established in the prior environment are reset as an adaptive measure. If an animal learns that a rustling sound is harmless in its familiar burrow, it would be evolutionary suicide to carry that complete inattention over to a novel foraging ground where predators lurk. The restoration of associability upon a context shift represents an adaptive reset of $\alpha$, demonstrating that the brain selectively gates attention in an ecologically sophisticated, context-dependent manner.
7. Intradimensional and Extradimensional Shift Paradigms
7.1 Theoretical Framework of Dimensional Attention
To demonstrate that selective attention operates across broader cognitive structures rather than merely tuning isolated, individual sensory cues, Nicholas Mackintosh turned to the theoretical framework of dimensional attention. Drawing upon the classic work of Donald Broadbent and the discrimination theories of Harry Harlow, Mackintosh posited that stimuli are not perceived as isolated, disconnected point-sources of energy. Instead, real-world stimuli are complex, multi-attribute compounds comprised of discrete, organized sensory dimensions, such as color, shape, spatial orientation, texture, and auditory pitch.
Mackintosh proposed the hypothesis of dimensional gating: when an organism discovers that a specific exemplar of a sensory dimension is a valid predictor of reinforcement, the attentional processing allocated to the entire dimension is elevated. If a pigeon learns that a red light predicts grain while a green light predicts non-reinforcement, the bird does not merely increase the associability parameter of the color red ($\alpha_{red}$). Rather, the entire cognitive channel dedicated to processing chromatic wavelengths undergoes an attentional upregulation ($\alpha_{\color}$). Simultaneously, uninformative sensory dimensions present during the task, such as geometric shape or pattern orientation, undergo dimensional suppression ($\alpha_{shape} \downarrow$).
This conceptualization fundamentally bridged animal conditioning with human comparative concept acquisition. It suggested that animal learning is not an uncoordinated collection of independent associative threads, but an organized process of dimensional set-formation. By tuning attention to entire physical dimensions, an organism acquires an abstract cognitive bias—an “attentional set”—that profoundly influences how it will approach completely novel sensory exemplars encountered in subsequent learning problems.
7.2 Experimental Validation via ID/ED Shift Paradigms
The definitive empirical test of dimensional attention lies in the classic Intradimensional (ID) versus Extradimensional (ED) shift paradigm. The methodological architecture of the ID/ED paradigm is one of the most sophisticated designs in behavioral neuroscience, structured to isolate acquired dimensional associability from specific stimulus-response habits.
The experiment unfolds across two primary stages:
- Stage 1 (Dimensional Acquisition): Animals are presented with compound stimuli varying along two distinct dimensions—for example, Dimension 1: Color (Exemplars: Red vs. Green) and Dimension 2: Shape (Exemplars: Circle vs. Triangle). For all subjects, Dimension 1 is relevant (Red is reinforced, Green is not), while Dimension 2 is irrelevant (Circles and Triangles are randomly distributed across trials). Animals train until they reach a high criterion of discriminative mastery.
- Stage 2 (The Transfer Shift): Crucially, all original stimulus exemplars are completely discarded and replaced with brand-new exemplars along the same dimensions. The new colors are Yellow and Blue; the new shapes are Square and Diamond. The subjects are divided into two critical experimental groups:
- Intradimensional Shift (ID) Group: The reinforcement contingency is governed by the same dimension that was predictive in Stage 1. Yellow (color) predicts reinforcement; Blue (color) predicts non-reinforcement. Shape remains irrelevant.
- Extradimensional Shift (ED) Group: The reinforcement contingency is reassigned to the dimension that was irrelevant in Stage 1. Square (shape) now predicts reinforcement; Diamond (shape) predicts non-reinforcement. Color is now irrelevant.
The empirical findings across thousands of replications are unequivocal: subjects in the ID shift group acquire the new discrimination with profound speed, mastering the problem in a fraction of the trials required by the ED shift group. This phenomenon is known as the ID shift superiority effect.
Purely elemental, non-attentional models of learning (such as Rescorla-Wagner) cannot predict ID shift superiority. Because the physical exemplars used in Stage 2 (Yellow, Blue, Square, Diamond) are completely novel, their associative values ($V$) are all zero, and their static physical saliences ($\alpha$) are unbiased. A non-attentional model must predict that learning a color discrimination should take the exact same number of trials as learning a shape discrimination. Mackintosh’s model explains the ID advantage perfectly: during Stage 1, the continuous predictive success of color exemplars drives the entire dimensional associability parameter ($\alpha_{\color}$) toward its maximum, while driving $\alpha_{shape}$ toward zero. When novel exemplars are introduced in Stage 2, the animal enters the problem with its color processing channel wide open and its shape processing channel actively suppressed. The ID group learns instantly because it is already attending to the correct informational dimension; the ED group struggles severely because it must first extinguish its dimensional attentional set and overcome the deep suppression of the shape dimension.
7.3 Species Differences in Dimensional Learning
The implementation of ID/ED shift paradigms across the phylogenetic tree provided Mackintosh and his contemporaries with profound insights into the evolution of cognitive flexibility. While the ID shift superiority effect is robustly observable across mammals and birds, comparative psychologists noted intriguing phylogenetic gradients in the magnitude, resilience, and operational characteristics of dimensional attention.
In non-human primates and human subjects, the formation of dimensional attentional sets is exceptionally rapid, often occurring within a handful of trials, and demonstrates powerful top-down cognitive control. Primates can execute multiple successive intradimensional shifts with near-zero transfer error, demonstrating an abstract cognitive rule: “attend to color.” Furthermore, when forced to execute an extradimensional shift, primates display a distinctive pattern of active hypothesis testing, rapidly abandoning an unrewarded dimensional set once its predictive validity breaks down.
In contrast, rodents (rats and mice) and avian species (pigeons) display dimensional transfer effects that are more tightly tied to continuous associative reinforcement. While rodents unquestionably exhibit robust ID shift superiority—demonstrating that their sensory systems are dynamically modulated by predictive relevance—their ability to rapidly disengage from a previously reinforced dimensional set during an ED shift is considerably more sluggish, characterized by persistent perseverative responding to the previously valid dimension. These comparative findings demonstrated that while Mackintoshian associability modulation represents an ancient, phylogenetically conserved associative mechanism common to all advanced vertebrates, the top-down cortical structures that govern the rapid, executive switching of dimensional gates evolved progressively greater flexibility in primate lineages.
8. Mackintosh (1975) Versus Rescorla-Wagner (1972): Comparative Analysis
8.1 Divergent Philosophies of Learning Regulation
The intellectual rivalry between the Rescorla-Wagner model (1972) and Mackintosh’s attentional theory (1975) represents one of the most profound theoretical debates in modern psychology. At the core of this dispute lies a fundamental disagreement over the cognitive locus of learning regulation: Is learning governed at the *output stage* through the processing of outcome discrepancy, or at the *input stage* through the selective gating of sensory cues?
The Rescorla-Wagner model embodies an outcome-processing philosophy. It assumes that the sensory input channels of an animal are wide open, passive, and indiscriminate. The nervous system registers every physical stimulus present in the environment with a fidelity fixed entirely by physical salience. The cognitive bottleneck occurs entirely at the unconditioned stimulus level: learning occurs if, and only if, the biological reinforcer surprises the animal. The critical parameter that regulates learning on any trial is the aggregate error term $(\lambda – \Sigma V)$. In this sense, the Rescorla-Wagner model is retroactive: the animal compares what it expected against what occurred, and updates all active associations in parallel based on that single outcome error.
In contrast, Mackintosh’s model embodies an input-processing philosophy. Mackintosh does not deny the reality of error-correction, but he insists that the primary regulatory control of learning occurs prior to associative integration, at the perceptual and attentional interface. The critical parameter is the dynamic stimulus associability vector, $\alpha_i$. Learning is not merely driven by whether the outcome was surprising; it is actively constrained by whether the animal was paying attention to the specific stimulus in the first place. Predictive utility acts as a forward-looking filter: past predictive success dictates current attentional bandwidth, which in turn determines the rate of future learning. Rescorla and Wagner model an animal that asks, “Was the world surprising today?”; Mackintosh models an animal that asks, “Which specific environmental features are worth my cognitive bandwidth?”
8.2 Critical Empirical Test Cases
To adjudicate between these two theoretical titans, experimental psychologists designed ingenious empirical test cases where the two models generated mutually contradictory behavioral predictions. Three specific paradigms stand out as decisive battlegrounds:
- Super-Conditioning (Conditioning to an Inhibitor): In this design, a target stimulus ($X$) is paired with a reinforcer in the presence of an established conditioned inhibitor ($A-$), a cue that reliably signals the absence of reinforcement ($V_A < 0$).
- Rescorla-Wagner Prediction: Because $V_A$ is negative, the compound associative sum is negative ($Sigma V = V_A + V_X < 0$). Consequently, the prediction error$(lambda – Sigma V)$ is massively enhanced—larger than if the US had occurred alone! The model predicts that stimulus $X$ will condition with extraordinary velocity, acquiring super-normal associative strength.
- Mackintosh Prediction: Stimulus $X$ is indeed the best available predictor of the reinforcer on those trials, so its associability $\alpha_X$ increases. However, Mackintosh correctly anticipated boundary constraints: if the overall situation is highly ambiguous or novel, associability adjustments proceed on a cue-by-cue relative validity metric. In multiple empirical variations, super-conditioning fails to occur at the extreme levels mandated by Rescorla-Wagner if the animal has already learned to discount the dimensional attributes of the compound.
- Recovery from Blocking via Extinction of the Blocking Stimulus: Following classic Kamin blocking ($A+$ followed by $AB+$), the blocked stimulus $B$ shows no conditioned response. What happens if an experimenter subsequently presents the blocking stimulus $A$ repeatedly in the absence of reinforcement ($A-$), driving $V_A$ down to zero?
- Rescorla-Wagner Prediction: The Rescorla-Wagner model predicts that the associative strength of stimulus $B$ ($V_B$) was frozen at zero during Stage 2 because the prediction error was zero. Extinguishing stimulus $A$ in isolation ($A-$) changes only $V_A$; it cannot retroactively alter $V_B$. Therefore, testing stimulus $B$ must still yield zero conditioned responding. (While modified comparator hypotheses later attempted to patch this, the core Rescorla-Wagner model is resolute).
- Mackintosh Prediction: Stimulus $B$ failed to condition because its associability $\alpha_B$ was degraded to zero. Extinguishing stimulus $A$ does not miraculously inject associative strength into $B$. However, Mackintosh’s model predicts that if compound $AB$ is subsequently reinforced, recovery of learning about $B$ will be profoundly retarded compared to a novel cue, because $\alpha_B$ remains at the floor. Mackintosh accurately predicted the enduring behavioral inertia of the blocked stimulus.
- The Associability of Partially Reinforced Stimuli: When a stimulus is paired with a reinforcer on an intermittent, probabilistic schedule (e.g., 50% partial reinforcement), how does its associability evolve?
- Rescorla-Wagner Prediction: The model has no mechanism for associability to change; $\alpha$ remains permanently fixed. The associative strength simply oscillates stably around a fractional value of $lambda$.
- Mackintosh Prediction: Because the partially reinforced stimulus generates persistent predictive errors on non-reinforced trials, it is a demonstrably poorer predictor than an alternative, continuously reinforced cue. Mackintosh predicts that animals will downregulate $\alpha$ for partially reinforced cues if a continuous predictor is available—a prediction confirmed in complex multi-cue relative validity settings, though challenged in isolated single-cue settings (which catalyzed the Pearce-Hall critique).
8.3 Strengths, Weaknesses, and Failures of Both Frameworks
A rigorous post-mortem of the Rescorla-Wagner and Mackintosh frameworks reveals that neither model represents an absolute, omnipotent account of associative learning. Each model possesses glaring structural blind spots that were exposed through decades of experimental replication.
The definitive triumph of the Mackintosh model is its natural, mathematically native explanation of pre-exposure phenomena. Latent inhibition, learned irrelevance, and intradimensional shift superiority flow directly from the core architecture of Mackintosh’s dynamic $\alpha$ update rule. The Rescorla-Wagner model is fundamentally paralyzed by these phenomena; it can only accommodate latent inhibition by bolting on ad-hoc, post-hoc auxiliary assumptions (such as inventing invisible contextual conditioning or arbitrarily declaring that unreinforced exposures modify peripheral stimulus thresholds).
Conversely, Mackintosh’s 1975 model exhibits notable vulnerabilities. Its greatest theoretical struggle lies in explaining sudden behavioral responses to dramatic unconditioned stimulus modifications. In Mackintosh’s system, if an animal is presented with an established predictor that suddenly delivers an unexpected, massive reinforcer, the cue was already the best available predictor ($|\lambda – V_A|$ was small), meaning its $\alpha$ was already high. However, if a cue is paired with an entirely unpredictable, erratic outcome, Mackintosh predicts that its associability should relentlessly decline because its predictive accuracy is low. Yet, empirical evidence frequently demonstrates the exact opposite: when an outcome becomes volatile and unpredictable, animals often dramatically increase their attention to the ambiguous cue to figure out the contingency!
This critical empirical anomaly demonstrated that Mackintosh’s formulation of attention was incomplete. Organisms do not merely pay attention to cues that are already established, highly accurate predictors (attention for action); they also desperately need to pay attention to cues whose consequences are uncertain, surprising, and unlearned (attention for learning). This realization precipitated the next great evolution in associative theory.
9. The Mackintosh Model Versus the Pearce-Hall (1980) Hybridization Debate
9.1 The Pearce-Hall Counter-Attentional Hypothesis
In 1980, John M. Pearce and Geoffrey Hall published a revolutionary paper in Psychological Review that directly inverted Nicholas Mackintosh’s central premise. While Mackintosh asserted that animals pay attention to cues that are accurate predictors of reinforcement, Pearce and Hall argued that organisms allocate attention primarily to stimuli whose consequences are uncertain, unpredictable, and surprising.
The operational logic of the Pearce-Hall model is rooted in computational efficiency. Why, Pearce and Hall asked, would an organism waste precious, metabolically expensive attentional bandwidth processing a stimulus whose outcome is already completely known and deterministic? If a tone is followed by food 100% of the time, the animal needs to execute the conditioned response automatically, but it does not need to waste processing power learning about the tone. Learning is only required when the nervous system does not know what will happen next. Therefore, Pearce and Hall proposed that the associability of a stimulus ($\gamma$) on trial $n$ is directly proportional to the absolute magnitude of the prediction error experienced on trial $n-1$:
$$\gamma_n propto |\lambda – \Sigma V|_{n-1}$$
Under the Pearce-Hall model, when an outcome is completely surprising ($|\lambda – \Sigma V|$ is large), stimulus associability surges to its maximum, opening the cognitive gates to allow rapid synaptic updates. As learning proceeds and the outcome becomes fully predicted, the prediction error approaches zero. Consequently, associability $\gamma$ decays toward zero. In complete opposition to Mackintosh, Pearce and Hall insist that established, infallible predictors suffer an attentional *decline*, becoming automated sensory-motor reflexes, whereas ambiguous or partially reinforced cues maintain perpetually elevated associability.
Pearce and Hall supported their model with striking empirical demonstrations. In one classic experiment, animals were trained on a stimulus that was followed by an unexpected change in reinforcement magnitude. Rather than ignoring the cue, the subjects displayed a massive, immediate enhancement in their rate of subsequent learning to that cue, proving that surprise directly reinvigorates stimulus associability.
9.2 Resolving the Apparent Paradox: Mackintosh vs. Pearce-Hall
For more than a decade, the field of associative learning was split between these two radically opposed, fiercely defended theories of attention. Did predictive validity increase attention (Mackintosh) or decrease it (Pearce-Hall)? The resolution to this intellectual deadlock arrived when researchers recognized that the two theories were fundamentally modeling two distinct forms of cognitive processing.
The modern consensus, pioneered by investigators such as Peter Holland, Peter Dayan, and Mike Le Pelley, distinguishes between “attention for action” (Mackintosh) and “attention for learning” (Pearce-Hall):
- Mackintoshian Attention (Attention for Action / Exploitation): When an animal must decide which behavioral response to execute in the present moment, it must rely on the most reliable, deterministic predictor available. The animal selectively focuses its behavioral, motor, and decision-making attention on cues with high predictive validity. This represents an exploitative cognitive state: relying on what has proven to work while ignoring distracting noise.
- Pearce-Hall Attention (Attention for Learning / Exploration): When an organism is updating its internal predictive models of the world, it must direct its plastic neural resources toward cues that are accompanied by surprise and predictive failure. If an established cue fails to deliver, the nervous system shifts into an exploratory, high-plasticity cognitive state, elevating the processing bandwidth of uncertain cues to determine the new environmental structure.
This functional dichotomy also maps cleanly onto the cognitive distinction between controlled processing and automaticity. When a cue is novel or its outcome unpredictable, it demands controlled, effortful processing (Pearce-Hall associability is high). Once the contingency is fully mastered, execution becomes automated, freeing up plastic cognitive resources (Pearce-Hall associability falls), while the cue maintains absolute behavioral dominance over alternative choices (Mackintoshian associability remains supreme).
9.3 Modern Integrated Systems: The Pearce-Mackintosh Syntheses
Recognizing the absolute ecological necessity of both attentional mechanisms, modern computational theorists formulated sophisticated hybrid associative systems that mathematically unify Mackintoshian validity-driven gating and Pearce-Hall surprise-driven plasticity. Prominent among these syntheses is the dual-attentional framework developed by Mike Le Pelley (2004) and computational neuroscientists like Peter Dayan.
In a unified dual-attentional model, the rate of learning about a conditioned stimulus is governed by the product of two distinct, dynamic associability parameters:
$$\Delta V_i = \alpha_i \times \gamma_i \times \beta (\lambda – \Sigma V)$$
In this integrated architecture, $\alpha_i$ represents the Mackintoshian parameter: it dynamically tracks the relative predictive validity of stimulus $i$ compared to competing environmental cues. Cues that are superior predictors maintain high $\alpha$ values, ensuring that they dominate behavioral control and are prioritized during sensory processing. Simultaneously, $\gamma_i$ represents the Pearce-Hall parameter: it tracks the absolute uncertainty or surprise associated with the outcome following stimulus $i$. If the outcome is fully predicted, $\gamma_i$ falls, dampening further plastic changes; if the outcome is erratic or surprising, $\gamma_i$ spikes, accelerating associative reassessment.
Empirical validation of these integrated dual-attentional mechanics has been achieved in modern human associative learning through high-speed eye-tracking technology. In human causal learning paradigms, researchers present participants with multi-element compound cues on computer screens and track their eye gaze fixations millisecond-by-millisecond while measuring their casual judgments. The eye-tracking data directly confirm the dual architecture: human participants consistently maintain their earliest and longest visual fixations on cues with high relative predictiveness (confirming Mackintosh’s attention for action), while showing rapid, reactive saccades to unpredicted, surprising cues following contingency shifts (confirming Pearce-Hall’s attention for learning).
10. Neurobiological Substrates and Brain Mechanisms of Attentional Conditioning
10.1 The Role of the Prefrontal Cortex in Dimensional Attention
The cognitive mentalism introduced by Nicholas Mackintosh found profound validation during the 1990s and 2000s as cognitive neuroscientists began mapping the physical, neural circuitry that implements selective associability adjustments. Chief among these neural structures is the prefrontal cortex (PFC), specifically the medial prefrontal cortex (mPFC) in rodents and the dorsolateral prefrontal cortex (dlPFC) and orbitofrontal cortex (OFC) in primates.
Neurotoxic lesion studies in rodents provide conclusive causal evidence that the prefrontal cortex is the biological engine driving dimensional attention. Rats with bilateral excitotoxic lesions to the prelimbic and infralimbic regions of the mPFC show completely intact basic classical conditioning: they learn simple CS-US associations at normal rates and exhibit normal motor responses. However, when these PFC-lesioned animals are tested in the classic Mackintosh Intradimensional/Extradimensional shift paradigm, their performance fractures. Lesioned animals can master the initial discrimination and successfully execute intradimensional shifts, but they become completely paralyzed when forced to execute an extradimensional shift. They perseverate obsessively on the previously predictive dimension, utterly unable to downregulate its attentional weighting or elevate the associability of the newly relevant sensory dimension.
Functional neuroimaging (fMRI) studies in human subjects during cue-competition tasks corroborate these animal lesion data. When human participants are exposed to compound stimuli where specific cues possess high relative predictive validity, blood-oxygen-level-dependent (BOLD) signals in the dorsolateral prefrontal cortex and anterior cingulate cortex (ACC) scale dynamically with the mathematical updates of Mackintosh’s $\alpha$ parameter. The prefrontal cortex actively sends top-down, glutamatergic feedback projections to early sensory cortices (such as visual area V4 or the primary auditory cortex), literally biasing sensory gain control: increasing neural firing rates in response to informative cues while suppressing neural representations of redundant sensory noise.
10.2 Subcortical and Neuromodulatory Systems
Beneath the prefrontal mantle, a deeply conserved network of subcortical nuclei and ascending neuromodulatory projections continuously computes, updates, and broadcasts the associability parameters demanded by the Mackintosh model. The most critical subcortical hub for attentional modulation is the cholinergic system of the basal forebrain, centered within the nucleus basalis magnocellularis (NBM) and the substantia innominata.
Acetylcholine (ACh) release in the sensory cortex and hippocampus acts as a physiological master-switch for cognitive plasticity and sensory associability. Seminal work by Peter Holland and Michela Gallagher demonstrated that neurotoxic lesions of cholinergic neurons projecting from the basal forebrain to the cortex selectively abolish changes in stimulus associability. Animals with depleted cortical acetylcholine can acquire baseline associations, but they fail to exhibit latent inhibition, fail to show learned irrelevance, and cannot downregulate attention to blocked stimuli. Acetylcholine release selectively enhances cortical processing of external sensory inputs while suppressing recurrent intrinsic connections, perfectly operationalizing the dynamic opening and closing of the Mackintoshian sensory gate.
Simultaneously, the mesolimbic dopaminergic system, originating in the ventral tegmental area (VTA) and projecting to the ventral striatum (nucleus accumbens), plays a nuanced, dual role. While Wolfram Schultz’s classical electrophysiology demonstrated that phasic dopamine firing encodes the Rescorla-Wagner reward prediction error ($(\lambda – \Sigma V)$), subsequent neurochemical assays revealed that tonic dopamine levels and localized dopamine transients in the striatum track the acquired salience and motivational associability of predictive cues. Stimuli that have acquired high Mackintoshian $\alpha$ values evoke immediate, localized dopamine release, signaling that these specific cues warrant behavioral priority.
Furthermore, the central nucleus of the amygdala (CeA) has been identified as an indispensable biological coordinator of attentional shifts. While the basolateral amygdala (BLA) encodes specific CS-US sensory representations, the CeA projects directly to the cholinergic basal forebrain. Disconnecting the basolateral amygdala from the central nucleus prevents the brain from updating stimulus associability when predictive contingencies shift, providing a precise neuroanatomical pathway linking the cognitive assessment of predictive error directly to the upregulation of cortical attention.
10.3 Hippocampal Involvement in Irrelevance and Gating
The hippocampal formation, alongside the adjacent parahippocampal and entorhinal cortices, forms the third critical biological substrate of Mackintosh’s attentional framework. The hippocampus has long been recognized as the primary neural locus for spatial mapping and episodic memory; however, in associative learning theory, the hippocampus serves as an essential informational sieve that computes environmental correlations and suppresses irrelevant stimuli.
Decades of neurobiological investigation have demonstrated that bilateral lesions to the hippocampus produce a dramatic, highly specific behavioral phenotype: the complete abolition of latent inhibition and learned irrelevance. If a rat sustains hippocampal damage and is subsequently pre-exposed to a tone for dozens of unreinforced trials, the animal completely fails to develop latent inhibition. When the tone is subsequently paired with a footshock, the hippocampally lesioned rat conditions with blinding speed, acquiring conditioned fear just as fast as an unexposed control subject. The animal acts as if it is entirely incapable of learning that a stimulus is irrelevant.
Electrophysiological recordings within the dentate gyrus and CA1 subfields of the hippocampus elucidate the computational mechanics of this gating system. When a novel stimulus is introduced, hippocampal pyramidal neurons display robust, synchronized firing. If that stimulus is repeated without biological consequence—rendering its predictive validity zero—hippocampal responses exhibit rapid, sustained habituation. The hippocampal-entorhinal circuit effectively constructs a statistical filter that identifies non-predictive, redundant sensory regularities in the environment. Through descending projections to subcortical neuromodulatory centers and the prefrontal cortex, the intact hippocampus actively instructs the sensory processing machinery to shut down attentional bandwidth to those cues. When the hippocampus is destroyed, this active inhibitory gating mechanism collapses, forcing the organism into a pathological state where every sensory perturbation, no matter how redundant, continues to be processed with maximum associability.
11. Contemporary Replications, Methodological Critiques, and Boundary Conditions
11.1 Replication Efforts in Modern Human Associative Conditioning
The transition of Nicholas Mackintosh’s theory from mid-century animal behavioral laboratories into 21st-century human cognitive science has yielded an extraordinary body of empirical literature. Modern researchers have adapted Mackintosh’s discriminative paradigms into sophisticated computerized causal learning tasks, allergy-prediction games, and financial market forecasting simulations designed to test human selective associability.
A typical human causal learning task presents participants with a fictional medical scenario: they play the role of a physician attempting to discover which food items cause allergic reactions in patients. Across sequential computer trials, participants view pictures of various food items (stimulus compounds) consumed by a patient, followed by the appearance or non-appearance of an allergic reaction (the outcome). Researchers meticulously arrange the food combinations to establish differential relative predictive validities across the food cues, precisely mirroring Mackintosh’s 1975 animal designs.
The empirical findings from these human paradigms have overwhelmingly replicated the core predictions of the Mackintosh model. When participants are introduced to a subsequent phase containing novel medical outcomes, they learn about previously predictive food items significantly faster than previously non-predictive items, even when the net causal strength of the items is meticulously controlled. Furthermore, by coupling these behavioral choice tasks with high-resolution eye-tracking, investigators such as Mike Le Pelley and Tom Beesley demonstrated that human gaze fixation duration serves as an overt, online physiological index of Mackintoshian associability. Human participants progressively extend their visual dwell time upon the most predictive cues in the visual array, literally refusing to look at cues that offer redundant informational utility. However, researchers have also identified distinct methodological divergences between human explicit casual judgment tasks and implicit autonomic conditioning (such as galvanic skin conductance responses), suggesting that while explicit judgments heavily recruit top-down prefrontal Mackintoshian gating, primitive autonomic responses sometimes remain vulnerable to low-level temporal contiguity.
11.2 Boundary Conditions and Experimental Anomalies
Despite its vast explanatory power, contemporary experimental psychology has rigorously demarcated the boundary conditions within which Nicholas Mackintosh’s attentional mechanics operate. Selective associability adjustments are not an invariant, universal feature of all conditioning; they are profoundly constrained by emotional state, arousal thresholds, and the biological nature of the reinforcer.
The most dramatic boundary condition emerges under conditions of extreme stress, life-threatening fear, and high-intensity trauma. While Mackintosh’s model operates flawlessly under mild-to-moderate appetitive and aversive conditioning schedules, it frequently breaks down when animals are exposed to severe, survival-threatening unconditioned stimuli, such as high-voltage electric shocks or intense predatory odor exposure. Under high fear states, the delicate, metabolically demanding prefrontal and hippocampal mechanisms that downregulate attention to redundant cues are overridden by the amygdala. In a high-intensity fear conditioning experiment, a redundant stimulus ($B$) presented in compound with an established fear predictor ($A$) often completely escapes blocking. The animal conditions vigorously to both cues, defaulting to a hyper-vigilant, non-selective associative strategy. From an evolutionary perspective, when death is on the line, the computational luxury of cognitive filtering is abandoned in favor of over-inclusive association.
A second critical boundary condition involves temporal contiguity overrides. Mackintosh’s model asserts that predictive validity trumps physical contiguity. However, temporal priority experiments show that if an uninformative, redundant cue is presented slightly earlier in time than a perfectly predictive cue (temporal precedence), animals frequently condition to the earlier, inferior predictor. The primitive temporal architecture of the subcortical nervous system places a massive premium on the earliest warning signal of an event, even if that signal is statistically noisy, occasionally overriding the slower, more sophisticated relative validity calculations executed by cortical networks.
11.3 Methodological Critiques of Mackintosh’s Experimental Logic
The intellectual rigor of the scientific method mandates that even the most celebrated theories be subjected to relentless methodological critique. Throughout the 1980s and 1990s, several prominent learning theorists, most notably Ralph Miller and Robert Rescorla himself, leveled pointed critiques against the empirical logic underlying Nicholas Mackintosh’s original 1975 experimental designs.
The primary critique focused on the pervasive challenge of performance factors masking associative values. In Mackintosh’s multi-stage transfer designs, the test phase measures overt behavioral responses (such as lever presses, magazine entries, or food cup pecks) to infer the underlying associability ($\alpha$) of the stimulus. Critics argued that the retarded acquisition observed to redundant or pre-exposed cues could be driven by non-attentional performance artifacts, such as:
- Conditioned Inattention / Motor Retardation: The animal may have learned a specific, overt competing motor behavior (e.g., looking away from the cue light) during the pre-exposure phase, which mechanically interferes with executing the conditioned response during the test phase, creating the illusion of a cognitive learning deficit.
- Floor and Ceiling Artifacts: Measuring the slope of acquisition curves in test phases is notoriously susceptible to measurement sensitivity. If baseline responding is too close to zero (floor) or rapid acquisition hits physiological maximums within three trials (ceiling), quantitative differences in learning rates can be wildly distorted, leading researchers to hallucinate variations in $\alpha$ when only subtle baseline differences in associative strength exist.
Furthermore, mathematical theorists questioned the uniqueness of Mackintosh’s operationalization of “relative predictive validity.” Mackintosh defined relative validity via the comparative inequality $| lambda – V_A | < | lambda - V_X |$. However, alternative mathematical formulations—such as calculating contingency ratios via delta-p statistics ($Delta P = P(US|CS) – P(US|neg CS)$) or Bayesian information gain metrics—can account for many of the same empirical findings without requiring the specific trial-by-trial update mechanics formulated in Mackintosh’s 1975 paper. These critiques did not demolish Mackintosh’s theory; rather, they catalyzed a generation of tighter experimental designs, culminating in the sophisticated modern eye-tracking and neural recording paradigms that ultimately vindicated Mackintosh’s core attentional thesis.
12. Lasting Legacy, Computational Extensions, and Modern Applications in Cognitive Science
12.1 Influence on Machine Learning and Neural Network Architectures
The intellectual footprint of Nicholas Mackintosh extends far beyond animal psychology, reaching deep into the computational foundations of modern artificial intelligence and machine learning. As contemporary computer scientists grapple with the challenges of training high-dimensional deep neural networks, they have repeatedly converged upon algorithmic mechanisms that are directly analogous to the acquired associability principles formulated by Mackintosh in 1975.
The most profound computational parallel exists between Mackintoshian associability and the revolutionary Self-Attention Mechanisms that power modern Transformer architectures (such as GPT-4, BERT, and modern large language models). In a Transformer network, the self-attention layer does not treat all tokens in a sequence equally, nor does it rely solely on their fixed positional contiguity. Instead, the model calculates dynamic, context-dependent “attention weights” based on the relative predictive utility between tokens. Tokens that provide the most informative, discriminative predictive context for generating the target representation receive heavily amplified attention weights, while irrelevant, redundant, or uninformative tokens are computationally masked out. This is nothing less than Mackintosh’s relative validity update rule running at digital scale across billions of parameters.
Furthermore, standard optimization algorithms in machine learning, such as AdaGrad, RMSprop, and the Adam optimizer, incorporate dynamic, parameter-specific learning rates. In traditional gradient descent, a single global learning rate ($eta$) is applied across the entire network—directly mimicking the static, non-plastic $\alpha$ assumption of the Rescorla-Wagner model. However, global learning rates lead to agonizingly slow convergence or catastrophic gradient explosions. Adaptive learning rate optimizers solve this by computing individual, dynamic learning rate vectors for each individual weight parameter, scaling the rate of learning inversely or directly with the historical reliability and variance of past gradients. By selectively accelerating learning on weights that receive sparse, highly informative gradients while dampening updates on noisy, redundant weights, modern machine learning implements the exact computational principle that Nicholas Mackintosh established when he freed $\alpha$ from physical invariance.
12.2 Clinical Implications for Psychopathology and Cognitive Disorders
In the domain of clinical psychiatry and abnormal psychology, Nicholas Mackintosh’s attentional theory of conditioning has served as a foundational diagnostic and explanatory framework for understanding the cognitive architecture of severe psychiatric illnesses, most notably schizophrenia.
For decades, psychiatrists observed that individuals suffering from schizophrenia display an inability to filter out irrelevant sensory stimuli, experiencing the world as an overwhelming, terrifying deluge of sensory information. In 1993, Philip Gray, Jeffrey Gray, and their colleagues demonstrated that the core cognitive deficit underlying acute schizophrenia is a complete, catastrophic failure of latent inhibition and learned irrelevance. When presented with standard latent inhibition paradigms, individuals with schizophrenia condition to pre-exposed, completely irrelevant stimuli just as rapidly as to novel stimuli. Their brain’s Mackintoshian sensory gate is permanently broken.
This failure maps directly onto modern neuropsychiatric concepts of aberrant salience. Due to hyperactive, dysregulated dopaminergic transmission in the mesolimbic pathway, the patient’s brain cannot compute zero associability. Environmental background noise, casual glances from strangers, or random street sounds are processed as carrying massive, vital predictive significance ($\alpha to 1$). The conscious mind, desperately struggling to make sense of why its subcortical neural hardware is assigning immense importance to trivial events, constructs elaborate, systematized paranoid delusions to rationalize this aberrant salience. Antipsychotic medications (dopamine D2 receptor antagonists) function therapeutically precisely by dampening this aberrant salience, restoring the brain’s ability to downregulate associability to irrelevant environmental stimuli.
Beyond schizophrenia, Mackintosh’s framework provides critical explanatory models for:
- Anxiety Disorders and Obsessive-Compulsive Disorder (OCD): In anxiety and OCD, individuals exhibit pathological hyper-associability toward minor, ambiguous sensory cues that remotely correlate with potential threat or contamination. The Mackintoshian gating system becomes rigidly locked, refusing to downregulate $\alpha$ even after thousands of non-reinforced exposures, leading to chronic behavioral avoidance and compulsive rituals.
- Substance Use Disorders and Addiction: In chronic addiction, environmental stimuli associated with drug consumption (e.g., drug paraphernalia, specific locations, social acquaintances) acquire extraordinarily high, enduring associabilities that resist extinction. These drug-paired cues command massive, involuntary attentional capture, triggering intense physiological cravings and behavioral relapse long after physical withdrawal symptoms have ceased.
12.3 The Enduring Epistemological Impact of Nicholas Mackintosh
When Nicholas Mackintosh passed away in 2015, he left behind a behavioral science transformed by his intellectual rigor, empirical craftsmanship, and theoretical vision. His enduring epistemological achievement was the brilliant, seamless synthesis of strict, unapologetic empirical behaviorism with the nuanced mentalism of cognitive science. Mackintosh demonstrated that one could study subjective, internal cognitive phenomena—such as attention, expectation, and informational utility—without sacrificing the uncompromising mathematical and methodological standards established by experimental animal psychology.
Mackintosh fundamentally dismantled the simplistic view of organisms as passive, empty associative vessels. He replaced the Pavlovian automaton with an active, informational organism: an agent that constantly interrogates its sensory horizon, computes the statistical reliability of its environment, allocates finite cognitive resources to what matters, and actively ignores what is redundant. In doing so, he provided the crucial intellectual bridge that carried psychology out of the rigid dogma of mid-century neo-behaviorism and propelled it into the modern era of cognitive neuroscience and computational modeling.
As associative learning theory continues to evolve in the 21st century—merging with Bayesian predictive processing frameworks, hierarchical reinforcement learning models, and optogenetic neural dissections—Nicholas Mackintosh’s 1975 attentional theory remains a bedrock pillar. His insight that learning is fundamentally gated by dynamic, acquired attention stands as an eternal truth of biological cognition: to learn about the world, an organism must first know where to look.
Conclusion
The trajectory of learning theory over the past century reveals a progressive awakening to the cognitive sophistication of the non-human and human mind. From the mechanistic reflexology of Pavlov and the rigid habit-formation equations of Hull, to the groundbreaking surprise-driven error correction of Rescorla and Wagner, psychology steadily chipped away at the passive assumptions of classical associationism. Yet, it was Nicholas Mackintosh who executed the decisive conceptual leap, proving that the mind does not merely react to the unexpected; it actively directs its own processing bandwidth.
By conceptualizing stimulus associability as a plastic, dynamic property dictated by relative predictive validity, Mackintosh unlocked the secrets of latent inhibition, learned irrelevance, intradimensional transfer, and selective cue competition. His model established that learning is an intensely competitive, highly structured process where environmental features must constantly prove their informational worth to breach the cognitive gates of the nervous system. Supported by exhaustive empirical research spanning multiple decades and confirmed across species from pigeons to primates, his principles now resonate through modern neurobiology, clinical psychopathology, and artificial intelligence.
Nicholas Mackintosh’s attentional theory of conditioning stands not merely as an alternative to the Rescorla-Wagner model, but as its essential, enduring intellectual partner. Together, these frameworks illustrate the magnificent computational balance of the biological brain: a system exquisitely tuned to act upon what is certain, learn from what is surprising, and relentlessly filter out the boundless noise of the sensory world.
References
- Broadbent, D. E. (1958). Perception and Communication. Pergamon Press. https://doi.org/10.1037/10037-000
- Dayan, P., Kakade, S., & Montague, P. R. (2000). Learning and selective attention. Nature Neuroscience, 3(Suppl), 1218–1223. https://doi.org/10.1038/81504
- Gray, J. A., Feldon, J., Rawlins, J. N. P., Hemsley, D. R., & Smith, A. D. (1991). The neuropsychology of schizophrenia. Behavioral and Brain Sciences, 14(1), 1–20. https://doi.org/10.1017/S0140525X00065055
- Holland, P. C. (1997). Brain mechanisms for changes in stimulus processing in associative learning. Philosophical Transactions of the Royal Society of London. Series B: Biological Sciences, 352(1363), 1997–2008. https://doi.org/10.1098/rstb.1997.0185
- Hull, C. L. (1943). Principles of Behavior: An Introduction to Behavior Theory. Appleton-Century-Crofts.
- Kamin, L. J. (1969). Predictability, surprise, attention, and conditioning. In B. A. Campbell & R. M. Church (Eds.), Punishment and Aversive Behavior (pp. 279–296). Appleton-Century-Crofts.
- Kruschke, J. K. (2001). Toward a unified model of attention in associative learning. Journal of Mathematical Psychology, 45(6), 812–863. https://doi.org/10.1006/jmps.2000.1354
- Le Pelley, M. E. (2004). The role of associative history in models of associative learning: A review and a hybrid model. The Quarterly Journal of Experimental Psychology Section B, 57(3), 193–243. https://doi.org/10.1080/02724990344000141
- Le Pelley, M. E., Beesley, T., & Griffiths, O. (2011). Overt attention and predictive learning: A review. Learning & Behavior, 39(2), 97–119. https://doi.org/10.3758/s13420-011-0024-8
- Lubow, R. E., & Moore, A. U. (1959). Latent inhibition: The effect of nonreinforced pre-exposure to the conditioned stimulus. Journal of Comparative and Physiological Psychology, 52(4), 415–419. https://doi.org/10.1037/h0046700
- Mackintosh, N. J. (1973). Stimulus selection: Learning to be ignorant. In R. A. Hinde & J. Stevenson-Hinde (Eds.), Constraints on Learning (pp. 75–96). Academic Press.
- Mackintosh, N. J. (1975). A theory of attention: Variations in the associability of stimuli with reinforcement. Psychological Review, 82(4), 276–298. https://doi.org/10.1037/h0076778
- Mackintosh, N. J. (1983). Conditioning and Associative Learning. Oxford University Press.
- Pavlov, I. P. (1927). Conditioned Reflexes: An Investigation of the Physiological Activity of the Cerebral Cortex (G. V. Anrep, Trans.). Oxford University Press.
- Pearce, J. M., & Hall, G. (1980). A model for Pavlovian learning: Variations in the effectiveness of conditioned but not of unconditioned stimuli. Psychological Review, 87(6), 532–552. https://doi.org/10.1037/0033-295X.87.6.532
- Rescorla, R. A., & Wagner, A. R. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In A. H. Black & W. F. Prokasy (Eds.), Classical Conditioning II: Current Research and Theory (pp. 64–99). Appleton-Century-Crofts.
- Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. Science, 275(5306), 1593–1599. https://doi.org/10.1126/science.275.5306.1593
- Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998–6008.
- Wagner, A. R., Logan, F. A., Haberlandt, K., & Price, T. (1968). Stimulus selection in animal discrimination learning. Journal of Experimental Psychology, 76(2p1), 171–180. https://doi.org/10.1037/h0025414