The quantification of voluntary action has long represented one of the most formidable challenges in the behavioral and cognitive sciences. Throughout the early decades of the twentieth century, the study of behavior was largely polarized between introspective mentalism and descriptive, qualitative reflexology. While early behaviorists demonstrated that environmental contingencies predictably alter the frequency of elicited and emitted responses, their analytical frameworks frequently operated under static, unichoice constraints. Organisms in natural ecological settings, however, do not encounter stimuli in isolation; instead, they are continually immersed in dynamic matrices of competing behavioral affordances, where every engaged activity necessarily precludes the execution of alternative repertoires.
The decisive epistemological breakthrough that transformed the experimental analysis of choice behavior occurred in 1961, when the American psychologist Richard J. Herrnstein formulated what is now universally recognized as the Matching Law. By deploying concurrent schedules of reinforcement within strictly controlled operant chambers, Herrnstein discovered an extraordinarily precise mathematical regularity: the relative rate of responding emitted toward a specific alternative matches the relative rate of reinforcement delivered by that alternative. This simple yet profound empirical theorem fundamentally altered the conceptual architecture of operant conditioning. It shifted the fundamental unit of behavioral analysis from absolute response counts to proportional resource allocation, establishing that choice is not an episodic, deliberate anomaly, but an ongoing, steady-state continuous process governed by relative environmental reinforcement density.
Over the past six decades, the Matching Law has evolved from an empirical generalization observed in avian operant chambers into a foundational pillar spanning behavioral economics, computational neurobiology, evolutionary ecology, and clinical psychology. Herrnstein’s initial linear formulation spurred the development of the Generalized Matching Law, quantitative formulations of Thorndike’s Law of Effect, melioration theory, and models of hyperbolic intertemporal discounting. Today, matching formulations provide the quantitative foundation for understanding diverse phenomena, ranging from synaptic plasticity in the basal ganglia and striatum to competitive tactical equilibria in professional athletics and the pathophysiology of substance use disorders. This treatise provides an exhaustive, mathematically rigorous, and empirically comprehensive examination of the Matching Law, tracing its historical genesis, mathematical formalization, neurobiological substrates, economic ramifications, and enduring scientific legacy.
1. Historical Context and the Genesis of the Matching Law
1.1 The Behaviorist Paradigm and Operant Conditioning Foundations
The genesis of quantitative choice modeling must be situated within the broader intellectual trajectory of twentieth-century radical behaviorism, particularly the experimental framework established by B. F. Skinner. Skinner’s introduction of the operant conditioning chamber provided a standardized methodology for isolating functional relationships between environmental stimuli, emitted motor actions, and consequential events. Central to Skinner’s epistemology was the adoption of response rate—measured via cumulative response recorders—as the primary dependent variable of behavioral science. By organizing reinforcement delivery across intermittent schedules based either on elapsed time (interval schedules) or completed response counts (ratio schedules), Skinner and his contemporaries demonstrated that the temporal distribution of reinforcement produces characteristic, highly reproducible steady-state patterns of behavior.
Despite these monumental empirical achievements, early operant psychology encountered significant theoretical limitations when attempting to extrapolate single-schedule dynamics to complex, naturalistic behavioral repertoires. In standard single-operant paradigms, an experimental subject is presented with a singular manipulandum, such as a lever or a response key. The animal faces a restricted binary condition: engage in the target response or engage in unmeasured, miscellaneous background behaviors (e.g., grooming, exploration, resting). Consequently, changes in absolute response rates observed under varying schedule parameters were frequently confounded by the uncontrolled, competing appeal of the background environment. This unichoice methodology could not systematically account for the relational nature of decision-making, wherein the value of an available outcome is inherently calibrated against the values of all currently accessible alternatives.
As radical behaviorism matured during the late 1950s, an epistemological shift emerged among quantitative researchers who sought to move the discipline beyond descriptive functional analyses toward axiomatic, predictive mathematical formulations. Modeling behavioral allocation under competing contingencies became an urgent scientific priority. To understand why an organism allocates time or energy to one activity over another, it was necessary to construct an experimental paradigm capable of presenting two or more explicitly programmed reinforcement contingencies simultaneously. This structural evolution required a formal departure from isolated, absolute rate evaluations toward the systematic quantification of relative response distributions, laying the groundwork for an authentic science of choice.
1.2 Richard Herrnstein’s Seminal 1961 Investigation
The definitive empirical realization of this quantitative transition arrived with Richard Herrnstein’s 1961 paper, titled “Relative and Absolute Strength of Response as a Function of Frequency of Reinforcement,” published in the Journal of the Experimental Analysis of Behavior. Operating within the Harvard Psychological Laboratories, Herrnstein devised a laboratory arrangement designed to isolate the fundamental laws governing concurrent behavioral allocation without the confounding artifacts typical of single-schedule paradigms.
Herrnstein utilized homing pigeons (Columba livia) maintained at reduced body weights to ensure stable motivational states. The experimental apparatus featured an operant chamber equipped with two horizontally adjacent response keys, each transilluminated by distinct visual stimuli. Crucially, Herrnstein deployed concurrent variable-interval schedules (abbreviated as conc VI VI). Unlike fixed-interval schedules, which produce cyclical scalloped response distributions, variable-interval schedules deliver reinforcers for the first response emitted after unsignaled, quasi-random intervals of time have elapsed, generating remarkably stable, continuous, and steady response rates across extended experimental sessions.
In this landmark experiment, Herrnstein systematically manipulated the relative frequency of food reinforcement allocated across the two response keys while holding the total reinforcement rate roughly constant. Across successive experimental conditions, the programmed schedules varied widely: in one condition, Key 1 might deliver forty reinforcers per hour while Key 2 delivered ten; in another condition, the contingencies were inverted, or balanced symmetrically at twenty-five reinforcers per hour on each manipulandum. Herrnstein allowed the subjects to engage with these concurrent schedules over tens of daily sessions until their behavioral outputs stabilized into invariant, steady-state equilibriums.
The resulting empirical data revealed an astonishing mathematical precision that stunned the behavioral neuroscience community. Rather than displaying an exclusive preference for the alternative that provided the higher reinforcement yield—an outcome predicted by traditional microeconomic theories of total utility maximization—the pigeons distributed their pecks across the two keys in near-perfect proportionality to the reinforcers obtained from each. When Key 1 yielded 70% of the total food presentations, the pigeons directed precisely 70% of their total pecks to Key 1. When the yield dropped to 30%, response allocation dropped in parallel to 30%. Herrnstein formalized this observation into an elegant empirical principle: the relative frequency of responding on a given operandum precisely matches its relative frequency of reinforcement.
1.3 Transition from Single-Schedule to Concurrent Schedules of Reinforcement
The conceptual transition from single-schedule paradigms to concurrent schedules of reinforcement marked an irreversible paradigm shift in behavioral science. In a single-schedule arrangement, the experimental investigator artificially isolates a response-reinforcer dyad, treating the organism as a closed system responding exclusively to the arbitrarily selected dependent variable. In contrast, concurrent schedules explicitly acknowledge the ecological reality of the organism: every behavioral niche is an arena of competing possibilities. In the natural world, a foraging organism never encounters a solitary food patch in a vacuum; the decision to exploit a given patch is continuously weighed against the temporal and energetic costs of abandoning that patch to forage elsewhere or engage in territorial defense, social courtship, or predator avoidance.
By conceptualizing all behavior through the lens of concurrent choice, Herrnstein demonstrated that absolute response rate is an epiphenomenon. The frequency with which an organism executes a specific action is fundamentally an artifact of the broader context of reinforcement—specifically, the ratio of reinforcement obtained from that action relative to the aggregate reinforcement obtained from all competing environmental alternatives. This realization dismantled the classical assumption that a reinforcer possesses an intrinsic, invariant strength capable of eliciting an absolute quantity of behavior regardless of external context.
Furthermore, the concurrent methodology illuminated the behavioral interdependence inherent in dual-operant systems. It established that an increase in reinforcement delivered for behavior A predictably and reliably suppresses the execution of behavior B, even if the absolute scheduled contingencies maintaining behavior B remain completely unaltered. This insight laid the conceptual groundwork for viewing all behavioral outputs as an ongoing, continuous allocation of a finite temporal budget among mutually exclusive alternatives. The experimental analysis of behavior was thereby elevated from an empirical cataloging of schedule response topographies to a predictive, quantitative science of choice allocation.
2. Mathematical Formulation of the Strict Matching Law
2.1 The Relative Ratio Equation
The mathematical formulation directly derived from Herrnstein’s 1961 empirical observations is known as the Strict Matching Law. In a canonical two-choice operant environment featuring concurrent schedules, let B1 and B2 denote the absolute response frequencies (or behavioral outputs) allocated to alternative 1 and alternative 2, respectively. Similarly, let R1 and R2 denote the absolute rates of reinforcement obtained from alternative 1 and alternative 2 over an identical observation interval. The Strict Matching Law states that the proportional behavioral output directed toward alternative 1 is equivalent to the proportional reinforcement obtained from alternative 1:
B1 / (B1 + B2) = R1 / (R1 + R2)
This formulation exhibits symmetrical properties across both alternatives. By algebraic rearrangement, an identical relationship holds for the alternative response pathway:
B2 / (B1 + B2) = R2 / (R1 + R2)
Alternatively, the strict matching relationship can be expressed in terms of odds or behavioral ratios rather than proportions. By dividing the proportional equation for alternative 1 by the proportional equation for alternative 2, the shared denominators cancel out entirely, yielding the ratio matching equation:
B1 / B2 = R1 / R2
This formulation is dimensionally consistent and geometrically linear. When relative response proportions, B1 / (B1 + B2), are plotted on the Cartesian ordinate against relative reinforcement proportions, R1 / (R1 + R2), on the abscissa, the theoretical strict matching relationship describes an invariant identity line passing through the origin with a slope equal to 1.0 and an intercept of 0.0. The mathematical elegance of this equation lies in its parameter-free nature: in its strict form, the model assumes that no systematic biological biases or structural distortions perturb the direct, proportional mapping of environmental reinforcement onto behavioral action.
2.2 The Absolute Response Rate Formulation and Hyperbolic Dynamics
While the relative ratio equation elegantly models choice between two programmed, concurrently monitored operants, it leaves unresolved the problem of predicting absolute response rates on an isolated, single schedule of reinforcement. In 1970, Herrnstein addressed this theoretical gap by publishing a landmark paper titled “On the Law of Effect,” in which he derived an absolute response rate formulation directly from the relational principles of matching.
Herrnstein posited that an organism in a single-operant setting is never truly exposed to a solitary schedule. Instead, the programmed experimental behavior (designated as B1) continuously competes with a nebulous reservoir of unmeasured, extraneous background activities (designated collectively as Be). These extraneous behaviors—which encompass grooming, preening, exploratory sniffing, wall-pecking, or resting—are maintained by an ambient pool of unprogrammed, extraneous environmental reinforcers (designated as Re). Applying the proportional matching logic to this total behavioral ecology yields:
B1 / (B1 + Be) = R1 / (R1 + Re)
Herrnstein introduced the physiological postulate that the total motor capacity of an organism per unit of time is finite and relatively invariant within standardized experimental contexts. If total behavioral output is defined as a constant asymptote, k, such that k = B1 + Be, substituting k into the equation yields Herrnstein’s famous hyperbolic equation for absolute response rates:
B1 = k · R1 / (R1 + Re)
This equation mathematically generates a negatively accelerating, hyperbolic response curve. When the programmed reinforcement rate R1 is zero, B1 is zero. As R1 increases relative to the extraneous reinforcement Re, response rate B1 ascends steeply. However, as R1 becomes exceedingly large relative to Re, the term R1 / (R1 + Re) asymptotically approaches unity, causing B1 to flatten out as it approaches the biological capacity limit k. The parameter Re possesses profound functional utility: it represents the half-saturation constant, precisely equal to the reinforcement rate required to elicit a response rate of half the maximal capacity (k / 2). This formulation provided the first rigorous quantitative bridge unifying single-operant dynamics with the relational mechanics of concurrent choice.
2.3 Quantification of Reinforcement Dimensions Beyond Frequency
In Herrnstein’s foundational paradigms, reinforcement was quantified solely along the single dimension of event frequency (number of reinforcement presentations per unit of time). However, reinforcers in both natural environments and sophisticated laboratory experiments inherently vary along several qualitative and quantitative physical dimensions. Subsequent investigations established that behavioral allocation is exquisitely sensitive to variations in reinforcement magnitude (amount or duration of access to a resource), reinforcement quality (nutritional, palatability, or hedonic characteristics), and reinforcement immediacy (the inverse of the temporal delay separating the response from reinforcer delivery).
To incorporate these critical variables, researchers generalized the strict matching framework into a multi-dimensional commodity model. If M represents the magnitude of the reinforcer (e.g., volume of sucrose solution, mass of food pellets, or duration of access to grain), D represents the temporal delay to delivery, and Q represents an empirically estimated qualitative index, the simple reinforcement term R can be replaced by an integrated, multidimensional composite value term, V:
V = f(R, M, Q, 1/D)
Assuming these reinforcement dimensions interact multiplicatively to establish total environmental payoff, the value of alternative i can be expressed as:
Vi = Ri · Mi · Qi · (1 / Di)
When substituted back into the fundamental matching ratio, behavioral allocation across two competing alternatives is dictated by the composite relative value equation:
B1 / (B1 + B2) = V1 / (V1 + V2)
Empirical verifications of this multi-attribute formulation—most notably conducted by William Baum, Howard Rachlin, and James Mazur—demonstrated that organisms trade off these variables with remarkable computational precision. For instance, an animal will match behaviorally between an alternative that yields a small reward delivered immediately and an alternative that yields a large reward delivered after an extended delay, providing a direct mathematical bridge between operant choice mechanics and the formal economic study of consumer choice and intertemporal evaluation.
3. Experimental Methodologies in Concurrent Reinforcement Schedules
3.1 Concurrent Variable-Interval Schedules (Conc VI-VI)
The successful empirical demonstration of the Matching Law requires precise methodological control over the experimental environment. Foremost among these methodological requirements is the careful selection of reinforcement schedules. The vast majority of quantitative choice experiments utilize concurrent variable-interval schedules (conc VI VI), rather than concurrent variable-ratio (VR) or fixed-ratio (FR) schedules. The rationale underlying this experimental convention is rooted deeply in the operational mechanics of interval versus ratio dependencies.
Under a variable-ratio schedule, reinforcement delivery is strictly contingent upon the completion of a predetermined, probabilistic number of emitted responses. In a concurrent ratio environment (e.g., conc VR 20 VR 40), the marginal probability of reinforcement per response is entirely independent of time and remains permanently higher on the leaner schedule (VR 20). Because every response allocated to the leaner schedule advances the subject closer to reinforcement faster than a response directed toward the richer schedule, any deviation from exclusive preference represents an immediate loss of utility. Consequently, concurrent ratio schedules inevitably produce an all-or-nothing behavioral dynamic known as “ratio-run lock-in,” wherein the organism exclusively directs 100% of its behavioral capacity to the schedule with the higher payoff probability, totally extinguishing responding on the alternate operandum.
In stark contrast, variable-interval schedules establish contingencies governed by the passage of time. Once the programmed interval on a given schedule has elapsed, a reinforcer is “set up” and remains held in biological storage until the subject emits a single operant response on that specific manipulandum. If an organism temporarily abandons a VI schedule to respond elsewhere, the abandoned schedule’s timer continues to run, steadily increasing the momentary probability that a reinforcer has become available on that neglected alternative. This temporal dynamic systematically penalizes exclusive preference; an animal that stubbornly remains on a rich VI 30-second schedule will miss out on the reinforcers quietly accumulating on a concurrent, leaner VI 90-second schedule. Therefore, concurrent VI-VI schedules maintain steady, continuous rates of responding across both operanda simultaneously, preserving behavioral sensitivity to shifting environmental payoffs and enabling the stable empirical measurement of choice allocation.
3.2 The Changeover Delay (COD) Procedure
Despite the functional advantages of concurrent variable-interval schedules, early experimentalists discovered an insidious methodological artifact that frequently obliterated the linear matching relationship: superstitious alternating chains. When an animal rapidly oscillates back and forth between two operanda (e.g., pecking Left, then Right, then Left), a reinforcer set up on the Right key might be delivered instantly upon the execution of a switch response. Under these conditions, the contiguous temporal pairing inadvertently reinforces the entire behavioral sequence: “Peck Left, switch to Right, peck Right.” Through the well-documented phenomenon of adventitious reinforcement, the animal develops a rapid, superstitious alternation pattern, behaving as though it were operating under a single complex chained schedule rather than two independent concurrent options. When this occurs, choice allocation decouples from the relative reinforcement rates, typically collapsing toward an indiscriminate 50:50 response distribution.
To eliminate this confounding artifact, Herrnstein engineered an essential procedural safeguard: the Changeover Delay (COD). The COD imposes an invariant temporal penalty—typically ranging between 1.5 to 3.0 seconds—immediately following any switch between operanda. When an organism leaves Key 1 and directs its first response to Key 2, the COD timer is initiated. During this brief interval, responses emitted on Key 2 are fully recorded by the apparatus but are fundamentally barred from producing reinforcement, even if the underlying VI schedule on Key 2 has already set up a reinforcer. Only responses emitted after the COD duration has fully expired can trigger the delivery of the reward.
The insertion of this minute temporal buffer completely severs the immediate temporal contiguity between the act of switching and the subsequent delivery of reinforcement. The COD forces the subject to persist at the newly selected alternative for a sustained moment before reaping the environmental harvest. Empirical investigations systematically varying COD parameters demonstrate that when the COD is absent (COD = 0 seconds), extreme undermatching or indiscriminate switching dominates the behavioral record. However, as the COD is raised to an optimal threshold (typically around 2.0 seconds in avian and rodent paradigms), stable, high-fidelity matching emerges. Conversely, if the COD is extended to excessively punitive durations (e.g., 20 or 30 seconds), it begins to function as an insurmountable travel cost, inducing overmatching and behavioral inertia.
3.3 Laboratory Paradigms and Operant Conditioning Chambers
The standard empirical apparatus for the investigation of matching dynamics is the computerized operant conditioning chamber, historically conceptualized as the “Skinner box.” In standard avian configurations, the experimental chamber is sound-attenuated, ventilated, and equipped with two or three horizontally aligned translucent response keys positioned at chest height on the intelligence panel. Each key is backed by an automated optical microswitch that records physical deflections requiring a precisely calibrated force (e.g., 0.15 Newtons). Multi-color LED projectors transilluminate the keys with specific visual wavelengths to establish discriminative control. Food delivery is governed by a solenoid-driven hopper that elevates a reservoir of grain into an illuminated aperture beneath the keys for precise durations (e.g., 3.0 seconds), during which the chamber’s overhead houselight is extinguished to isolate the reinforcing event.
In rodent paradigms, the keys are replaced either by low-inertia, retractable stainless-steel levers or illuminated nose-poke apertures equipped with infrared beam-break detectors. Infrared tracking matrices and automated force-transducer systems allow contemporary experimenters to record a continuous stream of behavioral metrics beyond simple binary response counts, including response latency, post-reinforcement pause duration, inter-response times (IRTs), switch frequency, and the specific physical topography of the operant action. Solid-state interfacing hardware coupled with real-time microprocessor control environments ensures that interval timers, schedule algorithms, and changeover delays are managed with millisecond temporal precision, completely eliminating human recording bias.
Crucially, the concurrent schedule methodology has transcended its initial non-human confines and been successfully translated into human laboratory settings. Human experimental protocols frequently employ computerized desktop tasks where participants allocate motor responses (such as mouse clicks, keyboard presses, or touchscreen touches) across concurrently displayed visual stimuli on a digital monitor. Reinforcers in human studies typically consist of point tallies exchangeable for real monetary payoffs, audible feedback tones, or consumable reinforcers. These translational human paradigms have confirmed that when adequate changeover penalties and stable steady-state conditions are enforced, humans distribute their behavioral investments in accordance with identical mathematical principles observed in non-human animals, demonstrating the profound phylogenic continuity of matching dynamics.
4. Deviations from Strict Matching: The Generalized Matching Law
4.1 Baum’s Power Function Formulation
Although Herrnstein’s Strict Matching Law provided a groundbreaking linear idealization of choice, subsequent decades of rigorous empirical research revealed that experimental data frequently exhibited systematic deviations from strict linear proportionality. Organisms often displayed behavioral allocations that were either insufficiently sensitive to differences in reinforcement yields, hypersensitive to richer alternatives, or skewed by structural motoric preferences. Recognizing that the parameter-free strict matching equation was too rigid to accommodate these universal behavioral nuances, William M. Baum proposed a transformative refinement in his classic 1974 paper, introducing what is now universally designated as the Generalized Matching Law.
Baum reformulated the matching relationship as a power function, incorporating two free empirical parameters: a sensitivity parameter (denoted as a) and an inherent bias parameter (denoted as b). Expressed in ratio form, the Generalized Matching Law states:
B1 / B2 = b · (R1 / R2)a
To facilitate empirical analysis and statistical parameter estimation via ordinary least squares regression, Baum applied a logarithmic transformation to both sides of the equation, yielding the celebrated log-ratio linear model of choice:
log(B1 / B2) = a · log(R1 / R2) + log(b)
In this transformed logarithmic space, the relationship between relative responding and relative reinforcement becomes strictly linear. The variable log(R1 / R2) serves as the independent variable on the abscissa, while log(B1 / B2) functions as the dependent variable on the ordinate. The exponent a represents the slope of the best-fitting linear regression line, indexing the behavioral system’s *sensitivity* to changes in relative reinforcement ratios. The parameter log(b) corresponds to the Cartesian y-intercept, quantifying any systematic, schedule-independent *bias* favoring one alternative over the other.
The Generalized Matching Law represents a massive methodological advance over the strict matching equation. When a = 1.0 and b = 1.0 (log(b) = 0), Baum’s formulation simplifies mathematically to Herrnstein’s original Strict Matching Law. However, by permitting a and b to vary freely as empirically determined parameters, the Generalized Matching Law can accurately fit, characterize, and isolate the distinct biological, cognitive, and procedural mechanisms responsible for deviations from perfect matching.
4.2 Mechanisms and Determinants of Undermatching
By far the most ubiquitous deviation observed in empirical choice literature is undermatching. Undermatching occurs whenever the sensitivity parameter falls systematically below unity (a < 1.0). In log-ratio space, this produces a regression slope that is flatter than the 45-degree line of strict matching. Behaviorally, undermatching signifies that the organism’s response allocation is less extreme than the relative reinforcement distribution; the subject allocates fewer responses to the richer schedule and more responses to the leaner schedule than predicted by strict matching.
A primary determinant of undermatching is inadequate stimulus discriminability. For an organism to match its behavioral output to relative environmental returns, it must possess the sensory and cognitive capacity to accurately perceive which alternative has delivered a given reinforcer, and to distinguish the current local payoff conditions. When the discriminative stimuli associated with the two schedules are visually or spatially ambiguous—or when schedules fluctuate rapidly before steady-state cognitive representations can form—the animal’s behavioral allocation regresses toward indifference (a 50:50 split), driving the sensitivity parameter a down toward zero.
Another prevalent mechanical cause of undermatching is the presence of an insufficient or absent Changeover Delay (COD). As discussed previously, when the temporal penalty separating switches between alternatives is set too low (e.g., COD < 1.0 second), the organism inadvertently experiences adventitious reinforcement of the switching response. This superstitious alternation creates an artificial floor on the frequency with which the leaner schedule is visited, diluting the measured behavioral divergence between the rich and lean options. Broad meta-analyses across hundreds of avian, rodent, and primate experiments confirm that the typical empirical value of a under well-controlled laboratory conditions hovers consistently around 0.80 to 0.90, reflecting the ubiquitous presence of perceptual noise, baseline environmental exploration, and minor transition artifacts in biological decision systems.
4.3 Overmatching and Heightened Sensitivity
The inverse deviation, known as overmatching, occurs whenever the sensitivity parameter significantly exceeds unity (a > 1.0). In graphical terms, overmatching manifests as a regression slope steeper than the theoretical 1.0 identity line. Phenomenologically, an overmatching organism displays an exaggerated, hypersensitive preference for the richer alternative: it over-allocates its behavior to the high-density schedule and severely under-allocates responses to the low-density schedule relative to the objective reinforcement ratio, frequently approaching absolute, near-exclusive preference.
Overmatching is almost universally generated by procedures that impose substantial transition costs, excessive physical friction, or extended travel times between the competing behavioral alternatives. When an experimenter implements an unusually severe Changeover Delay (e.g., COD > 10 seconds) or physical barriers that require the animal to traverse a physical corridor or hurdle an obstacle to access the alternate manipulandum, the energetic and temporal cost of switching becomes punitive. Under these high-cost conditions, the animal abandons continuous micro-sampling of the lean schedule. Instead, it exhibits behavioral inertia, remaining anchored to the richer option for prolonged runs to maximize immediate energy intake per switch.
In the natural world, overmatching corresponds to ecological patches separated by vast geographic expanses or perilous predatory terrain. When moving between foraging patches carries an acute risk of starvation or predation, natural selection strongly favors behavioral over-commitment to the local, known resource patch. Thus, overmatching does not represent a breakdown of operant rationality, but rather an adaptive response to high travel costs, where behavioral systems prioritize local exploitation over environmental exploration.
4.4 Asymmetry and Systematic Bias
The second free parameter in Baum’s generalized formulation, bias (represented as b in ratio form and log(b) in the log-ratio model), captures systematic asymmetries in choice allocation that persist irrespective of the programmed reinforcement schedules. A bias parameter of b = 1.0 (log(b) = 0.0) denotes a completely unbiased subject, where the regression line passes cleanly through the Cartesian origin. A positive or negative deviation in log(b) shifts the regression line vertically upward or downward, indicating an unconditional preference for one operandum over the other.
Systematic bias typically stems from three primary sources: physical asymmetries in the operant apparatus, unmeasured qualitative disparities in the reinforcer commodities, and inherent biological asymmetries within the organism itself. Physical apparatus asymmetries encompass factors such as subtle differences in the physical resistance of response keys, microswitch travel distances, background illumination levels, or spatial ergonomics (e.g., a lever positioned slightly closer to the food hopper). If Key 1 requires 0.10 Newtons of force to trigger while Key 2 requires 0.25 Newtons, the organism will display a persistent, immutable bias toward Key 1 across all reinforcement schedules, reflected as a positive log(b) intercept.
Similarly, when concurrent schedules deliver qualitatively distinct reinforcers—such as grain versus hemp seed, or condensed milk versus sucrose solution—the bias parameter log(b) directly isolates and quantifies the subject’s intrinsic hedonic preference for one commodity over the other, completely independent of their programmed delivery rates. In this capacity, Baum’s Generalized Matching Law serves as an exceptionally robust psychophysical measuring instrument. By systematically calculating the intercept shift across orthogonal schedule manipulations, behavioral scientists can quantitatively isolate the psychological value of reinforcer quality without confounding it with event frequency or schedule density.
5. Melioration Theory: The Proximate Mechanism of Choice
5.1 Conceptual Definition and Melioration Dynamics
Although the Matching Law brilliantly describes the macroscopic, steady-state distribution of behavior across concurrent schedules, it is inherently a *molar* empirical law; it describes aggregated relationships observed over extended temporal windows without specifying the real-time, dynamic decision rules executed by the organism at the *molecular* (moment-to-moment) level. To answer the mechanistic question of *how* organisms arrive at matching equilibrium, Richard Herrnstein and Drazen Prelec formulated Melioration Theory in 1991.
The term “melioration” is derived from the Latin meliorare, meaning “to make better.” Melioration theory posits a remarkably simple, continuous local optimization algorithm: an organism continuously redistributes its behavioral investments toward whichever alternative currently yields the higher *local* rate of reinforcement. The local rate of reinforcement on alternative i (denoted as ri) is defined not as the total reinforcers obtained divided by total session time, but as the reinforcers obtained from alternative i divided strictly by the time (or responses) allocated specifically to that alternative:
r1 = R1 / B1 and r2 = R2 / B2
The foundational behavioral dynamic of melioration dictates that if r1 > r2, the organism shifts behavioral allocation toward alternative 1, increasing B1 and decreasing B2. Conversely, if r2 > r1, behavioral allocation shifts toward alternative 2. This continuous behavioral reallocation naturally alters the experienced local reinforcement rates. As more behavior is directed toward alternative 1, its local yield eventually declines due to diminishing returns, while the local yield of the neglected alternative 2 rises as its uncollected rewards accumulate. The behavioral shift continues relentlessly until the system reaches a dynamic steady-state equilibrium where the local reinforcement rates across both alternatives are perfectly equalized:
R1 / B1 = R2 / B2
Crucially, simple algebraic cross-multiplication of this equalized local rate condition instantly reveals Herrnstein’s matching theorem:
B1 / B2 = R1 / R2
Melioration theory therefore provides the primary proximate behavioral engine driving matching: molar matching emerges as the inevitable, mathematical consequence of an organism continuously attempting to improve its immediate, local reinforcement return.
5.2 Melioration versus Global Maximization
The revelation that matching is driven by melioration provoked a fierce, foundational schism between radical behaviorism and classical neoclassical microeconomics. Classical economic rationality rests on the axiomatic assumption of *global maximization*: organisms possess complete, unhindered awareness of environmental production functions and invariably allocate behavior to maximize their total, aggregate utility (overall reinforcement rate) over an entire operational horizon.
Herrnstein and Prelec demonstrated mathematically and empirically that melioration and global maximization are fundamentally distinct behavioral processes that generate identical outcomes only under highly restricted, linear contingencies. Under standard concurrent variable-interval schedules, equalizing local rates of reinforcement coincidentally produces an overall reinforcement return that is extremely close to the mathematical maximum achievable under the schedule constraints. Because the performance curves around the maximization peak are exceptionally flat, early researchers erroneously assumed that subjects were calculating global utility optima.
However, when experimental contingencies are intentionally engineered such that local returns diverge systematically from global returns, organisms overwhelmingly follow the myopic path of melioration, consistently locking themselves into sub-optimal equilibrium states characterized by substantially reduced total reinforcement. In definitive experimental paradigms utilizing “contingent reinforcement schedules”—wherein the payoff schedules are dynamically linked such that allocating responses to the locally superior option actively drives down the future global reinforcement rate of both options—animals systematically fail to maximize total payoff. They persist in allocating behavior toward the locally richer alternative until local rates equalize, completely blind to the catastrophic collapse of their total environmental yield. The empirical triumph of melioration over global maximization conclusively proved that biological decision engines are governed by local, historical feedback loops rather than forward-looking, global optimization algorithms.
5.3 The Internalities Trap and Behavioral Traps
The divergence between melioration and global optimization reaches its most profound social and psychological expression in the concept of internalities. In classical economics, an “externality” occurs when an agent’s consumption decisions impose unpriced costs or benefits on other individuals. Herrnstein and Prelec coined the term *internality* to denote an intra-individual analogue: an internality occurs when an individual’s current consumption choice alters the future consumption possibilities or psychological payoffs experienced by that same individual at a later time.
When an activity generates negative internalities, choosing that activity produces an immediate local reward but simultaneously degrades the baseline yield of all future alternatives. This is the precise structural blueprint of an operant “behavioral trap,” most acutely exemplified by the pathology of chemical substance addiction. Consider an individual choosing between two behavioral streams: the consumption of an addictive substance (Alternative 1) and healthy, non-substance life activities (Alternative 2). Pharmacologically, Alternative 1 consistently delivers a higher instantaneous, local rate of reinforcement than Alternative 2 at any given moment of decision.
Driven by the biological mandate of melioration, the individual shifts time and behavioral resources toward drug consumption. However, sustained drug consumption generates severe negative internalities: physiological tolerance elevates the threshold for neurochemical reward, while the person’s professional, physical, and relational infrastructure deteriorates. Consequently, the baseline absolute value of *both* behavioral alternatives steadily drops. Yet, whenever the individual stands at the crossroad of choice, the local yield of the drug remains marginally superior to the severely degraded non-substance alternatives. Melioration compels the agent to continue consuming the drug, driving the individual down what Herrnstein called the “primrose path” toward a ruinous equilibrium characterized by catastrophic depression of aggregate life utility. Melioration theory thus provides an extraordinarily powerful, non-moralizing, purely operant etiology of human addiction and self-defeating behavior.
6. The Quantitative Law of Effect and Extraneous Reinforcement
6.1 Herrnstein’s Single-Schedule Formulation
Before the arrival of quantitative behaviorism, Edward Thorndike’s historic Law of Effect (1898) stood as an essentially qualitative axiom: behaviors followed by satisfying states of affairs are stamped in, whereas behaviors followed by annoying states are stamped out. For over seven decades, psychology lacked an axiomatic, mathematical equation expressing how varying densities of reinforcement transform into predictable levels of response output on an isolated schedule. Herrnstein achieved this breakthrough in 1970 by systematically applying concurrent matching mechanics to single-operant environments, establishing what is formally termed the Quantitative Law of Effect.
As established in Section 2.2, Herrnstein recognized that an organism in an operant chamber with access to a single lever is not living in an environmental vacuum. The programmed response B1 is simply one component of an omnipresent concurrent schedule where the alternate choice is the aggregate composite of all other available behaviors, Be, reinforced by extraneous background sources, Re. By substituting total capacity k into the matching proportional framework, Herrnstein yielded the canonical single-schedule equation:
B1 = k · R1 / (R1 + Re)
This formulation radically challenged the long-standing assumption that the relationship between reinforcement rate and response rate is a simple linear or monotonic power function. Herrnstein’s equation demonstrated that the relationship is fundamentally non-linear and hyperbolic. The response ceiling k dictates the physiological and motoric upper boundary of performance—the theoretical maximum response rate the animal could physically emit if all competing background reinforcers were completely eliminated (i.e., if Re = 0, then B1 = k regardless of R1).
The empirical validity of this hyperbolic formulation has been corroborated across thousands of experimental trials spanning diverse taxa, including fish, rodents, birds, dogs, non-human primates, and humans. Whether examining key pecking in pigeons, lever pressing in Sprague-Dawley rats, or academic task completion in neurodiverse children, empirical response curves map onto Herrnstein’s hyperbolic function with remarkable statistical fidelity, typically yielding coefficients of determination (R2) exceeding 0.95. Herrnstein’s formulation successfully operationalized Thorndike’s century-old intuition into a predictive, verifiable law of physical behavior.
6.2 Contextualizing Background Environmental Reinforcement (Re)
The introduction of the extraneous reinforcement parameter, Re, represents one of the most intellectually profound conceptual innovations in the history of behavior analysis. In classical psychology, if an investigator observed an unexplained plunge in an organism’s target response rate, the decline was routinely attributed to hypothetical internal psychological shifts, such as loss of motivation, fatigue, or cognitive boredom. Herrnstein’s formulation completely upended this internalist paradigm by demonstrating that target response rates can fluctuate wildly without any alteration whatsoever in the subject’s internal drive states or the programmed schedule of reinforcement, driven entirely by unmeasured fluctuations in the ambient environmental background, Re.
Extraneous reinforcement Re represents the aggregate sum of all reinforcing events maintaining all unmeasured alternative behaviors. In a standard rodent operant chamber, Re includes the sensory reinforcement derived from sniffing the chamber walls, tactile stimulation from self-grooming, kinesthetic feedback from rearing and stretching, and acoustic feedback from exploratory locomotion. If an experimenter unwittingly introduces an extraneous source of reinforcement into the chamber—for example, by leaving an olfactory trace of a conspecific, introducing an ambient draft, or using an apparatus with loose, rattling paneling—the magnitude of Re spikes dramatically.
Mathematically, examining the derivative of Herrnstein’s hyperbolic function with respect to Re demonstrates that as Re increases, the denominator (R1 + Re) expands, inevitably depressing the value of the quotient and driving down target responding B1. Conversely, if the background environment is stripped of all competing sensory affordances—rendering the chamber completely sterile and impoverished—Re approaches zero, forcing the animal to direct nearly all its motor capacity k into the solitary target response B1, even under exceedingly lean reinforcement schedules. Recognizing the active, continuous role of Re liberated behavioral science from viewing organisms as reactive automatons, positioning them instead as active ecological agents navigating an omnipresent field of competing reinforcement vectors.
6.3 Clinical Implications of Extraneous Reinforcers
The mathematical properties of the Quantitative Law of Effect carry monumental practical implications for clinical psychology, psychiatric rehabilitation, and Applied Behavior Analysis (ABA). Prior to the widespread adoption of quantitative matching formulations, clinical interventions designed to suppress aberrant or dangerous behaviors (e.g., severe self-injury, physical aggression, or destructive property destruction) relied heavily on aversive control mechanisms, including physical restraint, overcorrection, or targeted punishment schedules.
Herrnstein’s formulation revealed a revolutionary, non-punitive pathway for behavioral suppression. By conceptualizing the maladaptive target action as B1, maintained by an environmental schedule R1, clinicians can manipulate the client’s total behavioral allocation simply by inflating the background reinforcement density, Re. If the client’s surrounding environment is systematically enriched with dense, non-contingent access to potent alternative reinforcers—such as sensory integration items, highly preferred recreational activities, continuous social attention, and autonomous choices—the value of Re surges. As the denominator term (R1 + Re) expands exponentially, the absolute allocation of behavior to the maladaptive target action B1 collapses precipitously, without the need for aversive techniques.
This quantitative principle provides the mathematical foundation for modern non-contingent reinforcement (NCR) schedules and environmental enrichment paradigms in neurodevelopmental treatment centers and psychiatric facilities. Furthermore, it explains why behavior modification interventions implemented in highly sterile institutional settings frequently suffer catastrophic relapse when the individual returns to their natural home environment: a treatment environment artificially engineered to maintain an extremely low Re will cause target prosocial behaviors to appear artificially robust; once the individual transitions into an ecologically dense world teeming with competing extraneous reinforcers (elevated Re), the fragile target responses instantly crater. Long-term therapeutic sustainability requires calibrating intervention parameters against the mathematically estimated Re of the individual’s authentic ecological niche.
7. Neurobiological Correlates of Reinforcement Matching
7.1 Dopaminergic Value Coding and Reward Prediction Errors
For several decades following Herrnstein’s initial formulation, the Matching Law remained a functional, molar behavioral model lacking an identified neurobiological substrate. However, the rise of neuroeconomics and systems neuroscience in the late 1990s revealed that the brain’s internal decision architectures operate using computational algorithms that mirror the mathematical mechanics of matching. Central to this neurobiological convergence is the mesocorticolimbic dopamine system, originating in the ventral tegmental area (VTA) and substantia nigra pars compacta (SNc), and projecting broadly to the ventral striatum (nucleus accumbens) and prefrontal cortices.
Pioneering neurophysiological recordings conducted by Wolfram Schultz and colleagues established that the phasic firing of midbrain dopamine neurons encodes a quantitative metric known as the **Reward Prediction Error** (RPE). The RPE signal represents the mathematical delta between the subjective reward magnitude received (R) and the reward magnitude cognitively predicted to occur (V), formalized as:
RPE = Rreceived – Vpredicted
Computational neuroscientists, notably P. Read Montague, Peter Dayan, and Terry Sejnowski, demonstrated that this dopaminergic RPE is formally identical to the temporal difference learning algorithms utilized in computational reinforcement learning. In the context of concurrent schedules, phasic dopamine bursts do not simply register absolute sensory deliveries; they continuously adjust the synaptic weights linking discriminative sensory representations to motor commands in the striatum based entirely on relative outcome values.
When an animal engages with concurrent schedules, midbrain dopamine neurons dynamically calibrate their firing rates to reflect the relative reinforcement yields of the competing operanda. If a chosen alternative delivers a reinforcer whose value exceeds the ambient, background reward density of the chamber, a positive phasic dopamine burst occurs, inducing long-term potentiation (LTP) at corticostriatal synapses mediating that specific action. Conversely, if an alternative fails to deliver, or yields a reward inferior to the background expectation, dopamine firing dips transiently below baseline, triggering long-term depression (LTD). By continually updating the neurochemical balance between competing action representations, the phasic dopamine system provides the continuous, molecular value-tracking engine that drives steady-state behavioral matching.
7.2 Cortical and Striatal Circuits in Action Allocation
The systemic execution of matching behavior is coordinated across an integrated, multi-tiered neural network linking the prefrontal cortex, the posterior parietal cortex, and the basal ganglia. Within the prefrontal mantle, the orbitofrontal cortex (OFC) and the ventromedial prefrontal cortex (vmPFC) play an indispensable role in computing and updating the multi-attribute subjective values of competing commodities, integrating variables such as delay, effort, sensory quality, and current physiological state. Neurons within the OFC dynamically represent the anticipated value landscape, providing the essential value inputs required for comparative evaluation.
Simultaneously, the dorsolateral prefrontal cortex (dlPFC) and anterior cingulate cortex (ACC) track action-outcome associations, monitoring history-dependent choice stability and computing switching thresholds. When local reinforcement rates on a currently engaged operandum decline, the ACC registers heightened conflict and signals an increased likelihood of switching, effectively operationalizing the microscopic search mechanics of melioration.
Perhaps the most direct neurophysiological evidence linking single-neuron activity to Herrnstein’s equations was uncovered by Leo Sugrue, Greg Corrado, and William Newsome in a historic 2004 study published in Science. Recording from the lateral intraparietal area (LIP) of rhesus macaques performing a dynamic, concurrent visual foraging task governed by variable-interval schedules, the researchers observed that the electrophysiological firing rates of LIP neurons matched the matching law predictions with startling accuracy. The probability of the monkey executing a saccadic eye movement to a specific visual target was directly proportional to the relative fractional income obtained from that target over recent trials. LIP neurons were found to track the local, leaky temporal integration of reward income, demonstrating that sensorimotor association areas in the cerebral cortex explicitly construct neural representations of relative reinforcement ratios to guide competitive motor selection.
7.3 Computational Neurobiology and Synaptic Matching
To identify the precise biophysical mechanisms through which microscopic neural networks generate molar matching equations, theoretical neurobiologist H. Sebastian Seung formulated the Synaptic Theory of Matching in 2003. Seung addressed a profound theoretical paradox: how can individual neurons, which possess no global knowledge of aggregate reinforcement frequencies and are separated by vast spatial and temporal dimensions, collaboratively produce macroscopic behavioral matching?
Seung constructed a biophysical neural network model incorporating stochastic firing dynamics and plastic corticostriatal synapses governed by a reward-modulated covariance learning rule. In this computational framework, the synaptic weight (Wij) connecting a sensory input neuron j to a motor command neuron i is modified dynamically over time in accordance with the mathematical covariance between instantaneous neural activity and subsequent reward delivery:
ΔWij ∝ Cov(Actioni, Reward)
Seung demonstrated mathematically that when a neural network executes actions through a stochastic competitive process (such as a softmax decision rule or winner-take-all lateral inhibition) and updates its synaptic strengths via reward-modulated plasticity, the equilibrium state of the synaptic weights inevitably converges to the exact mathematical prediction of Herrnstein’s Matching Law. Synaptic connections mediating responses to Alternative 1 strengthen until the ratio of their transmission efficacy relative to Alternative 2 precisely mirrors the ratio of environmental rewards delivered across those two pathways.
This computational proof established that Herrnstein’s molar matching law is not merely an abstract phenomenological description, but the inevitable, emergent physical property of biological neural networks operating under localized, covariance-based synaptic learning rules. The Matching Law was thus conclusively anchored in the molecular architecture of the nervous system, unifying behavioral operant mechanics with the physical laws of neurobiology.
8. Behavioral Economics and the Matching Paradigm
8.1 Intertemporal Choice and Hyperbolic Discounting
The intersection of quantitative behavior analysis and economic theory catalyzed the birth of contemporary behavioral economics. One of the most monumental conceptual direct descendants of Herrnstein’s matching formulations is the modeling of intertemporal choice—the trade-offs organisms make between outcomes occurring at differing points in time. Classical economics historically modeled time preference through the lens of Paul Samuelson’s Exponential Discounted Utility model, which assumes that individuals discount future rewards at a constant, compounding interest rate per unit of time:
V = A · e–kD
Under exponential discounting, a person’s relative preference between two intertemporal rewards remains invariant over the passage of time, mathematically precluding the possibility of preference reversals. Real human and animal behavior, however, routinely violates exponential discounting, displaying profound dynamic inconsistencies, impulsivity, and regret.
In 1987, behavioral psychologist James Mazur demonstrated that intertemporal discounting is a direct mathematical derivative of Herrnstein’s Quantitative Law of Effect. By positioning reward immediacy (the reciprocal of delay, 1/D) as a fundamental dimension of reinforcement within the matching framework, Mazur formulated the famous Hyperbolic Discounting Equation:
V = A / (1 + k · D)
In this equation, V represents the present subjective value of the reward, A represents the nominal reward magnitude, D is the temporal delay to delivery, and k is an empirical discounting parameter indexing the individual’s degree of temporal impulsivity. Unlike the exponential curve, which decays at a constant percentage rate, the hyperbolic curve drops precipitously over immediate delays and flattens out over extended horizons.
This hyperbolic dynamic mathematically predicts the ubiquitous phenomenon of preference reversals. When both a Smaller-Sooner (SS) reward and a Larger-Later (LL) reward are separated from the chooser by substantial temporal distances, the individual rationally chooses the LL reward. However, as the passage of time draws the SS reward into the immediate temporal foreground, its hyperbolic value spikes dramatically due to the steep curvature of the function near D = 0. The subjective value of the immediate option surpasses that of the delayed option, provoking a sudden, irrational flip in choice toward the impulsive alternative. Mazur’s hyperbolic model—descended directly from Herrnstein’s operant equations—has become the standard theoretical framework for modeling impulsivity, executive dysfunction, financial mismanagement, and addictive behavior across economics, cognitive psychology, and clinical psychiatry.
8.2 Consumer Demand, Substitutability, and Complementarity
As behavioral economists expanded the matching framework into macro-behavioral settings, they incorporated foundational principles of microeconomic consumer demand theory, pioneered extensively by Steven Hursh. In natural economies, concurrently available commodities are rarely identical in their functional properties. The degree to which reinforcement matching holds depends heavily on the economic relationships of *substitutability*, *complementarity*, and *independence* between the available reinforcers.
Two commodities are designated as perfect substitutes if an organism can utilize either interchangeably to satisfy an identical biological drive (e.g., food pellets manufactured by differing commercial vendors, or two identical sucrose delivery apertures). Under conditions of perfect substitutability, concurrent choice data conform rigorously to the Generalized Matching Law. Increases in the price (ratio requirement) or reductions in the rate of one substitute trigger immediate, proportional reallocations of behavioral demand toward the alternative substitute, maintaining high sensitivity (a ≈ 1.0).
In stark contrast, when concurrent schedules deliver complementary commodities—goods that must be consumed jointly to generate utility, such as food and water—the mechanics of matching undergo profound structural alterations. If the price of food spikes precipitously under an operant schedule, the animal cannot simply abandon eating and divert its behavior exclusively to water consumption; doing so would result in fatal physiological distress. Under complementary conditions, demand becomes inelastic, and cross-price elasticity turns negative: an increase in the schedule requirement for food suppresses not only food-directed responding but also suppresses water-directed responding. Choice allocation under complementary constraints no longer tracks relative reinforcement rates in a simple linear fashion, yielding severe undermatching (a << 1.0).
Furthermore, behavioral economic experiments explicitly distinguish between closed economies (where the subject earns 100% of its total daily subsistence through experimental responding) and open economies (where the subject receives supplemental feeding outside the experimental session). In closed economies, organisms exhibit profound demand elasticity shifts, frequently driving sensitivity parameters toward extreme values to safeguard basic biological homeostasis. The integration of demand elasticity into the matching paradigm transformed behavior analysis from a simplistic study of response frequency into an advanced science of behavioral microeconomics.
8.3 Rationality and Economic Optimization Critiques
The empirical establishment of the Matching Law ignited a profound philosophical and theoretical critique of the foundational doctrine of “rational economic man” (Homo economicus). Neoclassical economics historically maintained that rational agents invariably execute choices that maximize an objective utility function. However, the ubiquitous empirical reality of matching—and its underlying proximate driver, melioration—demonstrated that biological organisms routinely and systematically violate global utility maximization.
This theoretical tension parallels the revolutionary work of Nobel laureate Herbert Simon on bounded rationality. Simon posited that biological decision-makers, facing severe computational limitations and incomplete information, do not execute global optimizations; instead, they utilize local heuristic decision rules that “satisfice.” Herrnstein’s Matching Law provides the precise mathematical embodiment of bounded rationality within operant systems. Melioration is a simple, computationally efficient, local sensory-motor heuristic: it requires no mental calculus of dynamic production functions or forward-looking integrals; it merely requires the organism to move toward whatever option is currently paying a higher local dividend.
Economists initially resisted this conclusion, attempting to construct complex mathematical proofs arguing that matching was simply an obscure, indirect manifestation of global maximization under unmeasured cognitive or temporal constraints. However, definitive laboratory experiments systematically refuted these claims. When researchers engineered schedule matrices where matching yielded catastrophic reductions in total earnings compared to global maximization strategies, animals and humans consistently drifted into matching equilibrium, fully forfeiting the maximum available utility. The Matching Law definitively demonstrated that the biological decision architecture of living organisms is fundamentally meliorative rather than globally optimizing, providing empirical verification for the foundational tenets of contemporary behavioral economics.
9. Human Social Interactions and Naturalistic Matching
9.1 Verbal Interaction and Conversational Dynamics
Although the Matching Law was derived through experimental analyses of avian operant behavior under rigid laboratory controls, its external validity was dramatically affirmed when applied to unstructured human social communication. The seminal investigation demonstrating naturalistic human matching was conducted in 1974 by Anthony Conger and Raymond Killeen, published in the Journal of the Experimental Analysis of Behavior.
Conger and Killeen recruited human participants to engage in dynamic, multi-person conversational discussions focused on social and moral issues. The participant was seated at a table across from two experimental confederates. The confederates delivered verbal reinforcers—consisting of positive social approvals such as head nods, agreement vocalizations (“mm-hmm,” “yeah”), and confirmatory praise—in accordance with hidden, independent, concurrent variable-interval schedules of reinforcement. Unknown to the participant, the experimenters systematically manipulated the relative frequency of social reinforcement distributed by Confederate 1 versus Confederate 2 over successive experimental blocks.
The primary dependent variable was the participant’s behavioral time allocation, quantified via hidden video recordings as the total duration of eye contact and continuous verbal engagement directed toward each confederate. The empirical results displayed an uncanny correspondence to Herrnstein’s equation: human participants matched the relative proportion of their verbal communication time to the relative proportion of social approvals delivered by each confederate with mathematical precision. Subsequent replications expanded these findings across naturalistic group therapy sessions, elementary school classroom discussions, and family dyad interactions, confirming that human social dynamics—long assumed to be governed entirely by complex mentalistic agency—are fundamentally anchored in quantitative operant matching mechanics.
9.2 Athletic Competition and Tactical Play Selection
Outside the clinical and conversational laboratory, naturalistic athletic competitions represent an exceptionally rich, high-stakes testing ground for the Matching Law. Elite competitive sports feature experienced human performers operating under extreme motivational states in environments where choices, consequences, and tactical distributions are rigorously and comprehensively recorded by statistical databases.
A classic application of matching theory to sports analytics was pioneered by Timothy Vollmer and colleagues, who analyzed shot selection in professional basketball (NBA). Basketball players continuously face a concurrent choice between two offensive behavioral streams: attempting a two-point field goal (a higher-probability event located closer to the basket) versus attempting a three-point field goal (a lower-probability event located beyond the arc). Applying Baum’s Generalized Matching Law, researchers modeled the relative allocation of shot attempts as a function of the relative point yield obtained from each category. The resulting analyses demonstrated that professional basketball teams distribute their offensive shot attempts across the two-point and three-point options in near-perfect accordance with the Generalized Matching Law, exhibiting sensitivity values (a) hovering remarkably close to 1.0.
Similar quantitative matching equilibria have been documented across professional American football (NFL) play-calling distributions (allocating offensive snaps between rushing plays and passing plays) and soccer penalty kicks (shot direction vs. goalkeeper dive direction). The underlying mechanism driving matching in sports is directly linked to defensive tactical adjustments: if an offensive team over-utilizes the run, the opposing defense adjusts its personnel to counter the run, driving down its local efficiency. The offense is forced to meliorate, redistributing plays until the local yields of running and passing reach dynamic parity. The Matching Law thus provides a descriptive and predictive quantitative model for competitive tactical equilibrium in elite human athletics.
9.3 Digital Media and Online Engagement Dynamics
In the contemporary digital era, the architecture of the internet, social media platforms, and mobile computing interfaces operates as an omnipresent, highly engineered concurrent schedule of reinforcement. Software engineers and user-interface (UI/UX) designers deliberately construct digital environments to capture and maintain human behavioral allocation, utilizing variable schedules of notification delivery, content loading, and algorithmic reward feedback.
When an individual navigates a smartphone, their behavioral allocation—quantified as continuous screen time, dwell time, and click-through rates (CTR)—is distributed across concurrent digital behavioral streams (e.g., social media feeds, video streaming platforms, digital messaging apps, and productivity tools). Platforms like TikTok, Instagram, and X (formerly Twitter) utilize algorithmic recommendation systems that act as continuous, high-density variable-interval schedules. The user’s operant response (the vertical swipe or infinite scroll) is intermittently reinforced by the presentation of highly stimulating, hedonic content, interspersed unpredictably with mundane or neutral posts.
Applying the Generalized Matching Law to digital media consumption demonstrates that users allocate screen time across competing applications in direct proportion to the relative density and immediacy of algorithmic reinforcement provided by those applications. Furthermore, because mobile applications deploy instantaneous push notifications and zero-latency transition architectures, the Changeover Delay between apps is virtually nonexistent. As predicted by matching mechanics, the absence of a COD drives rapid, compulsive task-switching, continuous distraction, and severe undermatching. Digital consumers frequently find themselves ensnared in severe melioration traps: the immediate local yield of scrolling an algorithmic feed consistently outcompetes the delayed, effortful yields of occupational work, physical exercise, or sleep, locking millions of individuals into suboptimal equilibria characterized by widespread digital fatigue and fractured attention.
10. Applications in Clinical Psychology and Applied Behavior Analysis
10.1 Functional Assessment and Problem Behavior Etiology
In the clinical domain of Applied Behavior Analysis (ABA), the Matching Law has revolutionized the functional assessment and conceptualization of severe problem behaviors. Aberrant actions—including self-injurious behavior (SIB), physical aggression directed toward caregivers, severe tantrums, and property destruction—were historically viewed as manifestations of internal neurochemical imbalances or unmanageable psychiatric defects. Quantitative behavior analysis transformed this view by proving that severe problem behavior is an orderly, operant choice governed by concurrent reinforcement schedules operating within the individual’s living environment.
In natural clinical settings, an individual diagnosed with intellectual disabilities or autism spectrum disorder continuously chooses between emitting an aberrant response (B1) and emitting an adaptive, communicative response (B2). The reinforcers maintaining these behaviors typically encompass social attention (smiles, verbal soothing, eye contact), tangible items (access to preferred foods or toys), sensory stimulation, or immediate escape from demanding academic or vocational tasks (negative reinforcement). Clinical research reveals that problem behavior frequently emerges and stabilizes because the natural environment inadvertently programs a dense, immediate, high-magnitude schedule of reinforcement for aberrant actions, while prosocial communication is met with lean, highly delayed reinforcement.
A child who must emit forty polite verbal requests before a caregiver responds (a lean VI schedule) may discover that banging their head against the floor instantly commands the immediate, focused physical attention of multiple adults (a rich, dense VI schedule with zero delay and massive magnitude). Under the mathematical dictates of the Matching Law, the child’s behavioral allocation shifts overwhelmingly toward self-injury. Through functional behavior assessments, clinical analysts do not merely identify the functional reinforcer maintaining a problem behavior; they quantitatively measure the competing schedules operating across the entire behavioral ecology, identifying the structural reinforcement imbalances that compel the client to allocate behavior to destructive topographies.
10.2 Differential Reinforcement of Alternative Behaviors (DRA)
The foundational clinical treatment derived from the matching paradigm is Differential Reinforcement of Alternative Behaviors (DRA). The mechanical objective of a DRA protocol is to systematically alter the environmental contingency matrix such that the individual voluntarily reallocates behavior away from the maladaptive response (B1) and toward an adaptive, replacement response (B2, such as functional communication training, sign language, or icon exchange).
Historically, early behavioral modification attempts frequently failed because clinicians simply placed the problem behavior on extinction (zero reinforcement) without establishing an adequate replacement schedule. Under extinction alone, the individual often displays severe extinction bursts, emotional aggression, and eventual treatment failure. Guided by the Matching Law, modern DRA engineers interventions along all physical dimensions of the composite value equation: rate, immediacy, magnitude, and quality. To guarantee the decisive reallocation of behavior to the prosocial alternative, the treatment schedule is engineered such that:
VDRA >> Vproblem
The clinician ensures that the communicative replacement behavior is reinforced on a continuous, dense, immediate schedule using maximum-magnitude rewards, while the problem behavior is placed on an extinction schedule (or an exceedingly lean, delayed, low-magnitude contingency). As the relative value term for the adaptive response ascends, the matching equation dictates that the proportional allocation of behavior to the problem response inevitably drops toward zero.
Furthermore, quantitative matching provides the foundational theoretical framework for understanding and preventing clinical treatment relapse, specifically the phenomena of *resurgence*, *renewal*, and *reinstatement*. Resurgence occurs when an alternative behavior maintained by DRA is subjected to schedule thinning or accidental extinction; under these conditions, the matching equation demonstrates that the individual must allocate their motor capacity k somewhere, inevitably driving behavior back toward the historical problem behavior. By systematically accounting for extraneous reinforcement variables (Re) and meticulously pacing schedule thinning, applied behavior analysts can engineer long-term, durable behavioral resilience that resists relapse when clients transition into complex, uncontrolled community environments.
10.3 Addiction and Substance Use Disorders
The application of the Matching Law to substance use disorders represents one of the greatest triumphs of quantitative behavioral science. Rather than viewing addiction as an unalterable biological disease state, the behavioral economic paradigm conceptualizes substance use as an extreme operant choice allocation. Pharmacological substances of abuse—such as cocaine, methamphetamine, heroin, alcohol, and nicotine—deliver immediate, high-magnitude, hyper-potent chemical reinforcers directly to the brain’s mesolimbic dopamine reward circuitry. When mapped against the matching equation, the local value of drug consumption (Vdrug) is colossal.
Critically, the pathology of addiction is severely exacerbated by environmental poverty and social isolation. When an individual’s living environment contains minimal accessible sources of non-drug reinforcement—characterized by chronic unemployment, social disenfranchisement, untreated psychiatric trauma, and lack of recreational infrastructure—the extraneous reinforcement pool (Re) or the value of alternative behaviors (Valt) collapses toward zero. In the mathematical formulation of matching, as the denominator term representing alternative reinforcement shrivels, the proportional allocation of behavior directed toward the drug pathway asymptotically approaches 1.0 (exclusive preference), regardless of the horrific long-term downstream costs.
This quantitative realization directly gave birth to Contingency Management (CM), developed extensively by Stephen Higgins and colleagues. Contingency Management is one of the most empirically robust, scientifically validated psychosocial interventions for substance use disorders in existence. CM protocols deliberately introduce potent, highly structured alternative economic reinforcers into the user’s environment. Clients who provide verified drug-negative urine samples receive immediate, tangible monetary vouchers exchangeable for goods, services, job training, and healthy recreational activities.
By systematically injecting dense, dependable alternative reinforcement into the client’s operant ecology, CM dramatically elevates Ralt. As the relative reinforcement ratio shifts, the matching equation mathematically demands a reallocation of behavioral investment away from drug-seeking behavior and toward the voucher-producing recovery lifestyle. Long-term longitudinal studies confirm that clients who establish rich, durable networks of non-drug environmental reinforcers maintain long-term abstinence, proving that substance use disorders are profoundly sensitive to the quantitative reallocation of relative environmental contingencies.
11. Theoretical Debates and Competing Behavioral Models
11.1 Molecular Maximization versus Molar Matching
Throughout the history of quantitative behaviorism, no dispute has been more intensely contested than the molar-molecular debate. This fundamental epistemological controversy centers on the temporal grain of behavioral control: is matching an authentic, primary biological law that operates over broad, aggregated time horizons (the molar perspective), or is it merely a statistical artifact generated by fine-grained, momentary choices aimed at maximizing short-term payoff (the molecular perspective)?
The molar perspective, championed by Richard Herrnstein, William Baum, and Howard Rachlin, argues that behavior cannot be decomposed into isolated, discrete atomistic reflexes. They maintain that the true functional units of behavior are extended temporal streams of activity distributed across ongoing environmental schedules. According to molar theorists, organisms track aggregate reinforcement correlations over broad time windows, directly matching their macroscopic behavioral distributions to macroscopic reinforcement densities without executing microscopic calculations.
Conversely, the molecular optimization camp—spearheaded by researchers such as Charles Shimp, Allen Neuringer, and John Nevin—asserts that behavioral science must ground itself exclusively in immediate, momentary events. Shimp formulated mathematical models demonstrating that organisms continually compute the *momentary probability of reinforcement* for each available alternative. At any given split second, an animal simply executes the single motor action that has the highest conditional probability of being reinforced right now. Shimp proved mathematically that when animals execute momentary maximization strategies under concurrent variable-interval schedules, their aggregated behavioral output across a session will coincidentally mimic the matching law with exceptional precision. While both paradigms can account for standard steady-state choice, subsequent experiments utilizing non-linear, dynamic feedback schedules designed to explicitly dissociate molecular predictions from molar predictions have generally shown that choice behavior is governed by local, short-term molecular feedback loops, supporting the mechanics of melioration over abstract, long-range molar integration.
11.2 Optimal Foraging Theory and Biological Adaptation
While experimental psychologists were uncovering the matching law within operant chambers, evolutionary biologists and behavioral ecologists were independently investigating the mechanics of animal decision-making in the wild, establishing the discipline of Optimal Foraging Theory (OFT). Central to OFT is Eric Charnov’s celebrated Marginal Value Theorem (MVT), which predicts how long a foraging animal should remain within an exploiting food patch before abandoning it to search for another, given the energetic costs of travel and the ambient patch density of the overall habitat.
The convergence between the Matching Law and Optimal Foraging Theory represents one of the most beautiful syntheses in modern behavioral biology. In nature, food resources are distributed across distinct ecological patches that experience progressive resource depletion as the organism feeds. An animal foraging in a patch experiences diminishing marginal returns over time, a dynamic functionally identical to an operant variable-interval schedule. If an animal were to execute strict, inflexible matching (a = 1.0) under all conditions, it would continually distribute its foraging time across all patches, including dangerously lean or distant locations. However, natural selection operates on net reproductive fitness, which frequently demands overmatching (specialization on high-yield patches) when travel costs are perilous, or undermatching (sampling across multiple patches) when environments are highly volatile and unpredictable.
From an evolutionary vantage point, the Generalized Matching Law’s sensitivity parameter (a) represents an adaptive calibration mechanism for resolving the ubiquitous **exploration versus exploitation trade-off**. Undermatching (a < 1.0) is not an error or biological flaw; it is an evolutionarily conserved risk-management strategy. In natural ecosystems, environmental yields are never static; a rich berry patch may be devoured by competitors or desiccated by drought, while a lean patch may suddenly flourish. By maintaining a baseline undermatching allocation, the foraging animal continually expends a small fraction of its behavioral budget sampling suboptimal patches, ensuring it detects environmental shifts and preserves survival resilience. The Matching Law is thus revealed as the foundational computational architecture through which biological organisms optimize ecological patch exploitation.
11.3 Cognitive and Computational Decision Models
In contemporary cognitive science and computational neuroscience, choice behavior is rarely modeled using static algebraic formulas. Instead, researchers deploy dynamic, stochastic process models, prominent among which are **Sequential Sampling Models** and the Drift-Diffusion Model (DDM). These computational models represent decision-making as the continuous accumulation of noisy sensory or mnemonic evidence over time toward an internal decision boundary.
In a standard two-alternative forced-choice framework, evidence favoring Alternative 1 drives an internal decision variable upward, while evidence favoring Alternative 2 drives it downward. The rate of evidence accumulation—the *drift rate* (ν)—is dictated by the relative subjective value of the competing stimuli. Once the decision variable strikes the upper or lower absorption threshold, the corresponding motor action is executed. Computational modeling demonstrates that when the drift rate of a diffusion process is continuously modulated by the relative reinforcement rates experienced across concurrent schedules, the resulting steady-state choice distributions naturally converge to Baum’s Generalized Matching Law.
Similarly, within the architecture of modern artificial intelligence, algorithmic agents operating under reinforcement learning frameworks (such as Q-learning, SARSA, or Actor-Critic architectures) routinely exhibit matching dynamics. In an Actor-Critic architecture, the “Critic” evaluates incoming reward prediction errors to update the value representations of environmental states, while the “Actor” updates a parameterized policy matrix governing action selection probabilities. When an artificial agent equipped with an Actor-Critic algorithm is placed within a simulated concurrent variable-interval environment, its behavioral allocation across extended episodes settles precisely into the matching law. This remarkable cross-disciplinary convergence proves that the Matching Law is an invariant mathematical equilibrium point for any adaptive learning system—biological or silicon—that adjusts action selection probabilities based on experiential reward feedback.
12. Methodological Challenges, Synthesis, and Future Directions
12.1 Measurement Difficulties and Analytical Complexities
Despite its conceptual elegance, the empirical investigation of the Matching Law presents complex methodological, analytical, and statistical hurdles that require rigorous experimental control. Foremost among these challenges is the mathematical treatment of zero-response bins. In Baum’s log-ratio formulation:
log(B1 / B2) = a · log(R1 / R2) + log(b)
If an experimental subject emits zero responses on an operandum during an observation block (B2 = 0), or if a schedule delivers zero reinforcers over an interval (R2 = 0), the resulting ratios involve division by zero, generating undefined values that cannot be transformed logarithmically. Researchers historically addressed this by artificially adding arbitrary constants (e.g., adding 0.5 or 1.0 to all response and reinforcement counts), but this arbitrary practice introduces severe mathematical distortions and artificially flattens the estimated regression slope, generating spurious undermatching artifacts. Contemporary quantitative analysts resolve this by deploying non-linear mixed-effects models, zero-inflated Poisson regressions, and hierarchical Bayesian estimation techniques that fit raw, untransformed count data directly to the power function without requiring logarithmic linearization.
A second major analytical complication is the presence of high autocorrelation and non-stationarity in longitudinal choice time series. Successive responses emitted by an organism in an operant chamber are rarely independent; they exhibit profound sequential dependencies, post-reinforcement pausing patterns, and local micro-clustering. Applying standard Ordinary Least Squares (OLS) regression to autocorrelated data violates fundamental Gauss-Markov assumptions, drastically underestimating standard errors and inflating the statistical significance of parameter estimates. Modern choice analysis resolves this by implementing autoregressive moving average (ARIMA) error structures, generalized estimating equations (GEE), and dynamic state-space filtering algorithms capable of isolating true, steady-state sensitivity parameters from transient, history-dependent fluctuations.
12.2 Expansion to Multi-Choice and Networked Environments
For the first four decades of its history, the Matching Law was evaluated almost exclusively within binary choice architectures (N = 2). However, real-world ecosystems and human societal networks present agents with hundreds of concurrently available alternatives (N >> 2). Expanding the mathematical architecture of the Generalized Matching Law to accommodate multi-choice decision spaces represents a vibrant frontier of contemporary behavioral science.
In an N-choice environment, the proportional matching formulation for any given alternative i generalizes to:
Bi / ∑ Bj = (Ri)ai / ∑ (Rj)aj
However, fitting multi-choice empirical data reveals complex, non-linear cross-commodity interactions that cannot be captured by simple pairwise comparisons. The presence of an irrelevant third alternative can radically alter the relative choice ratio between two original alternatives, violating the microeconomic axiom of the Independence of Irrelevant Alternatives (IIA)—a phenomenon known in cognitive psychology as the “decoy effect” or “asymmetric dominance.”
Furthermore, researchers are increasingly scaling matching formulations up to analyze complex, decentralized social networks. In multi-agent digital environments, such as decentralized financial markets, crypto-economic incentive structures, and viral social communication networks, hundreds of human or automated agents act simultaneously as concurrent schedules of reinforcement for one another. Utilizing the tools of statistical physics, graph theory, and evolutionary game theory, computational social scientists are discovering that collective behavioral migrations, crowd panics, and the viral propagation of digital memes follow macro-level matching dynamics. The Matching Law is thus emerging as a critical quantitative bridge linking micro-level operant learning rules to macro-level sociological and economic network phenomena.
12.3 The Lasting Legacy of Richard Herrnstein’s Behavioral Science
The scientific legacy of Richard J. Herrnstein stands as one of the most enduring and transformative monuments in twentieth-century behavioral biology. When Herrnstein published his modest empirical observations on pigeon key-pecking in 1961, behaviorism was dominated by qualitative assertions and mechanistic stimulus-response models that were increasingly vulnerable to the burgeoning cognitive revolution. By formulating the Matching Law, Herrnstein demonstrated that voluntary, emitted operant behavior could be modeled with the mathematical rigor, predictive precision, and physical universality characteristic of the natural sciences.
Herrnstein’s quantitative formulations executed an epistemological revolution across multiple disciplines. In theoretical psychology, he dismantled the concept of absolute reinforcement, establishing that all behavior is choice, and that every isolated action is fundamentally an allocation derived from relational environmental matrices. His subsequent extensions—the absolute Quantitative Law of Effect, melioration theory, and hyperbolic discounting models developed with his colleagues—provided the foundational mathematical scaffolding that enabled the birth of contemporary behavioral economics and neuroeconomics.
Today, the Matching Law remains as vibrant, relevant, and scientifically generative as it was over six decades ago. It guides the neuroscientist recording dopamine kinetics in the striatum, the applied behavior analyst designing life-saving interventions for self-injurious children, the quantitative sports analyst optimizing tactical point allocations, and the computer scientist engineering adaptive reinforcement learning architectures in artificial intelligence. By discovering the simple, elegant mathematical ratio that maps environmental consequences onto physical action, Richard Herrnstein unlocked a fundamental natural law of biological life: a universal equation that governs the continuous, dynamic dance of choice across all living things.
Conclusion
The Matching Law represents a triumphant synthesis of empirical methodology, mathematical precision, and ecological validity. From its origins in the Harvard operant conditioning laboratories to its contemporary integration with systems neuroscience, behavioral economics, and clinical psychology, Herrnstein’s formulation has weathered more than six decades of rigorous empirical scrutiny. It has demonstrated that choice is not an inscrutable, metaphysical mystery, but an orderly, deterministic physical process governed by the relative distribution of environmental payoffs.
Through its successive theoretical evolutions—from the Strict Matching Law to Baum’s Generalized Matching Law, the absolute Quantitative Law of Effect, and Herrnstein and Prelec’s Melioration Theory—the paradigm has continually expanded its explanatory power. It accounts not only for optimal steady-state choice distributions, but also for the cognitive noise of undermatching, the structural rigidities of overmatching, the behavioral traps of addiction, and the dynamic temporal reversals of intertemporal choice. By seamlessly linking the molecular firing rates of midbrain dopamine neurons to the macroscopic dynamics of human social systems, the Matching Law stands as one of the most profound, unifying theoretical achievements in the history of the behavioral sciences.
References
- Baum, W. M. (1974). On two types of deviation from the matching law: Bias and undermatching. Journal of the Experimental Analysis of Behavior, 22(1), 231–242. https://doi.org/10.1901/jeab.1974.22-231
- Baum, W. M. (1979). Matching, undermatching, and overmatching in studies of choice. Journal of the Experimental Analysis of Behavior, 32(2), 269–281. https://doi.org/10.1901/jeab.1979.32-269
- Charnov, E. L. (1976). Optimal foraging, the marginal value theorem. Theoretical Population Biology, 9(2), 129–136. https://doi.org/10.1016/0040-5809(76)90040-X
- Conger, R., & Killeen, P. (1974). Use of concurrent operants in analyzing talking behavior in humans. Journal of the Experimental Analysis of Behavior, 21(2), 273–284. https://doi.org/10.1901/jeab.1974.21-273
- Herrnstein, R. J. (1961). Relative and absolute strength of response as a function of frequency of reinforcement. Journal of the Experimental Analysis of Behavior, 4(3), 267–272. https://doi.org/10.1901/jeab.1961.4-267
- Herrnstein, R. J. (1970). On the law of effect. Journal of the Experimental Analysis of Behavior, 13(2), 243–266. https://doi.org/10.1901/jeab.1970.13-243
- Herrnstein, R. J., & Prelec, D. (1991). Melioration: A theory of distributed choice. Journal of Economic Perspectives, 5(3), 137–156. https://doi.org/10.1257/jep.5.3.137
- Higgins, S. T., Alessi, S. M., & Dantona, R. L. (2002). Voucher-based incentives: A substance abuse treatment innovation. Addictive Behaviors, 27(6), 887–910. https://doi.org/10.1016/S0306-4603(02)00295-4
- Hursh, S. R. (1980). Economic concepts for the analysis of behavior. Journal of the Experimental Analysis of Behavior, 34(2), 219–238. https://doi.org/10.1901/jeab.1980.34-219
- Mazur, J. E. (1987). An adjusting procedure for studying delayed reinforcement. In M. L. Commons, J. E. Mazur, J. A. Nevin, & H. Rachlin (Eds.), Quantitative Analyses of Behavior: Vol. 5. The Effect of Delay and of Intervening Events on Reinforcement Value (pp. 55–73). Lawrence Erlbaum Associates.
- Montague, P. R., Dayan, P., & Sejnowski, T. J. (1996). A framework for mesencephalic dopamine systems based on predictive Hebbian learning. Journal of Neuroscience, 16(5), 1936–1947. https://doi.org/10.1523/JNEUROSCI.16-05-01936.1996
- Rachlin, H. (1971). On the tautology of the matching law. Journal of the Experimental Analysis of Behavior, 15(2), 249–251. https://doi.org/10.1901/jeab.1971.15-249
- Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. Science, 275(5306), 1593–1599. https://doi.org/10.1126/science.275.5306.1593
- Seung, H. S. (2003). Learning in spiking neural networks by reinforcement of stochastic synaptic transmission. Neuron, 40(6), 1063–1073. https://doi.org/10.1016/S0896-6273(03)00761-X
- Shimp, C. P. (1966). Probabilistically reinforced choice behavior in pigeons. Journal of the Experimental Analysis of Behavior, 9(4), 443–455. https://doi.org/10.1901/jeab.1966.9-443
- Simon, H. A. (1955). A behavioral model of rational choice. The Quarterly Journal of Economics, 69(1), 99–118. https://doi.org/10.2307/1884852
- Skinner, B. F. (1938). The Behavior of Organisms: An Experimental Analysis. Appleton-Century-Crofts.
- Sugrue, L. P., Corrado, G. S., & Newsome, W. T. (2004). Matching behavior and the representation of value in the parietal cortex. Science, 304(5678), 1782–1787. https://doi.org/10.1126/science.1094765
- Vollmer, T. R., & Bourret, J. (2000). Application of the matching law to evaluate the allocation of two- and three-point shots by college basketball players. Journal of Applied Behavior Analysis, 33(2), 137–150. https://doi.org/10.1901/jaba.2000.33-137