The study of operant conditioning underwent a profound epistemic transformation during the middle of the twentieth century. For several decades following the publication of B.F. Skinner’s foundational monographs, the experimental analysis of behavior had been fundamentally qualitative and descriptive, characterized by the demonstration of functional control over single response topographies under isolated reinforcement schedules. Operant chambers systematically evaluated cumulative response curves, documenting how distinct schedules of reinforcement—fixed, variable, ratio, and interval—sculpted the temporal patterning of an organism’s emitted actions. While these early paradigms definitively established that behavioral rates could be systematically modulated by environmental contingencies, they remained constrained by a critical ecological limitation: non-human animals and humans rarely, if ever, navigate an environment offering only a solitary course of action in total isolation from competing alternatives.
In 1961, Harvard psychologist Richard J. Herrnstein published a seminal investigation entitled “Relative and Absolute Strength of Response as a Function of Frequency of Reinforcement” in the Journal of the Experimental Analysis of Behavior. Herrnstein sought to move behavior analysis beyond the isolated operant baseline by confronting experimental subjects with simultaneous, continuously available, yet independent alternatives. By arranging concurrent variable-interval schedules of reinforcement for pigeons pecking visual keys, Herrnstein introduced an elegant quantitative methodology for studying choice in a continuous, free-operant framework. What emerged from this experimental architecture was an empirical relationship of striking regularity: organisms distributed their behavioral output across competing response keys in exact proportionality to the distribution of reinforcers delivered by those keys. This empirical phenomenon, christened the Matching Law, provided the very first mathematical invariant for operant choice behavior.
The implications of Herrnstein’s matching law radiated far beyond the mechanics of avian key-pecking. By demonstrating that behavioral allocation across multiple responses obeys rigorous mathematical regularities, Herrnstein established the intellectual foundation for quantitative behavior analysis, modern behavioral economics, and neuroeconomics. The matching relationship fundamentally unified rate and choice, revealing that all behavior is ultimately an allocation problem occurring against a dynamic background of competing reinforcement sources. The following comprehensive monograph details the historical antecedents, precise methodological mechanics, mathematical formulations, theoretical debates, and clinical applications that constitute the rich lineage of Herrnstein’s matching law experiment.
1. Historical Antecedents and the Emergence of Quantitative Behavior Analysis
1.1 The Transition from Qualitative Operant Conditioning to Quantitative Models
During the nascent period of radical behaviorism, B.F. Skinner established the primary baseline of operant psychology around the concept of response rate. Utilizing the cumulative recorder, early behavioral researchers tracked the frequency with which a laboratory animal—typically a rat or a pigeon—actuated a mechanical manipulandum, such as a lever or a pecking disc. In these early frameworks, response rate served as the primary dependent variable, conceptualized as a direct index of habit strength, habit strength probability, or operant response probability. Schedules of reinforcement, cataloged extensively by Ferster and Skinner in their classic 1957 compendium, demonstrated that response patterns were fundamentally shaped by environmental contingencies rather than internal drive states. However, these experimental designs were largely qualitative and idiographic, relying heavily on visual inspection of cumulative records rather than predictive algebraic equations.
As the discipline evolved into the late 1950s, experimental psychologists encountered severe methodological and theoretical ceilings within the single-operant paradigm. In a standard single-key chamber, an animal is presented with a binary condition: respond on the available manipulandum or engage in unspecified, unmeasured “other” behaviors, such as grooming, resting, or exploring the perimeter of the chamber. Because these alternative behaviors were neither tracked nor programmed, the experimental architecture could not quantitatively model complex decision-making, preference, or economic tradeoffs. If an animal decreased its response rate on a single lever, the experimenter could not determine whether this drop reflected satiation, fatigue, or the transient intrusion of a more attractive competing reinforcer. The single-operant arrangement artificially stripped the environment of its natural complexity, where organisms continuously navigate multiple competing schedules of reward.
Simultaneously, the broader landscape of experimental psychology was undergoing a profound quantitative revolution. Mathematical modeling was gaining unprecedented traction through the pioneering work of theorists such as William K. Estes with stimulus sampling theory, Robert Bush and Frederick Mosteller with mathematical models for learning, and R. Duncan Luce with choice axioms. Psychophysics had long demonstrated that sensory modalities conformed to strict mathematical power laws, as articulated in S.S. Stevens’ psychophysical scaling formulations. Operant investigators began to recognize that if behavior analysis were to attain parity with the natural and physical sciences, it needed to transcend descriptive taxonomies of schedules and identify robust, predictive mathematical invariants that could describe behavioral phenotypes across varying motivational, environmental, and phylogenetic conditions.
This philosophical evolution spurred behavioral researchers to transition from searching for local response rates to formulating generalized mathematical laws of allocation. Rather than viewing an organism’s behavior as an absolute quantity triggered by isolated stimuli, the new quantitative behavior analysis began conceptualizing behavior as a continuous, dynamic stream. In this emerging paradigm, every emitted action represented a selection made at the expense of an infinite array of competing behavioral possibilities. The central challenge lay in engineering an experimental apparatus and a theoretical apparatus sophisticated enough to measure how organisms systematically distribute their limited behavioral budgets when faced with concurrent, independently operating reinforcement alternatives.
1.2 Richard Herrnstein’s Academic Background and Theoretical Influences
Richard J. Herrnstein entered Harvard University’s Department of Psychology during this transformative period, pursuing his doctoral studies under the direct mentorship of B.F. Skinner. Herrnstein brought to the Harvard laboratory an analytical temperament that synthesized Skinnerian radical behaviorism with rigorous quantitative methodology. Early in his academic career, Herrnstein was deeply immersed in problems of sensory discrimination, stimulus control, and psychophysical measurement in non-human animals. His initial experimental publications tackled complex psychophysical questions in pigeons, demonstrating that non-verbal organisms could perform remarkably refined sensory judgments when appropriate operant contingencies were instituted. This psychophysical training instilled in Herrnstein a persistent drive to discover lawful, quantitative scaling properties governing the relationship between external inputs—such as reinforcers—and behavioral outputs.
Beyond the Skinnerian milieu, Herrnstein was acutely aware of developments in neoclassical microeconomics, particularly mathematical utility theory and consumer choice models. Microeconomic models presupposed that human economic agents distributed financial resources across market commodities to maximize subjective utility, governed by marginal rates of substitution. While standard behavioral psychologists of the late 1950s often viewed economic theories with suspicion due to their reliance on mentalistic assumptions like utility and cognitive expectation, Herrnstein recognized that microeconomic questions concerning resource allocation bore an unmistakable formal similarity to the ways an organism allocates its finite physiological time and physical energy across environmental opportunities.
Herrnstein’s fundamental conceptual leap was the operationalization of response distribution as an objective, observable index of subjective value. If an organism were confronted with two distinct sources of reinforcement that varied in frequency, magnitude, or quality, its internal valuation of those alternatives could not be measured by peering into a cognitive black box; instead, it would be manifested directly in the proportion of physical behavior directed toward each alternative. By bridging Skinner’s free-operant methodology with the formal rigor of choice theory, Herrnstein sought to strip utility of its subjective, mentalistic connotations, translating it into an empirically measurable behavioral distribution. This theoretical synthesis set the stage for Herrnstein’s departure from single-operant baselines toward concurrent schedule architectures.
1.3 The Paradigmatic Shift toward Choice and Preference Assessment
Prior to the establishment of the matching law, the prevailing methodology for investigating animal choice was the discrete-trial apparatus, epitomized by the traditional T-maze, Y-maze, or Wisconsin General Test Apparatus (WGTA). In a standard T-maze trial, an animal is placed at the base of a runway, navigates to a choice point, selects either the left or right arm, consumes whatever reinforcer is present, and is then physically removed by the experimenter. While the T-maze yielded valuable historical insights into spatial learning and habit strength, it suffered from severe methodological artificiality. Each discrete trial imposed an arbitrary beginning, middle, and end upon behavior, enforced extensive experimenter handling, and obliterated the continuous, self-paced flow that characterizes natural behavioral repertoires. Discrete trials measured only the terminal outcome of choice—a binary percentage of left versus right turns—completely obscuring the temporal dynamics and rate-dependent properties of decision-making.
Herrnstein recognized that choice in the natural world is rarely composed of segmented, isolated trials orchestrated by an external agent. In ecological niches, foraging animals navigate environments where multiple food patches, predatory threats, and mating opportunities exist concurrently and persistently over time. An animal does not execute an instantaneous, irrevocable choice between patches; rather, it continuously allocates portions of its daily time budget, transitioning back and forth among various foraging sites as resources deplete and replenish. To capture this ecological reality in the laboratory, behavioral science required a continuous, free-operant arrangement where an organism could switch fluidly between alternatives at any self-determined moment without experimenter intrusion.
The implementation of free-operant concurrent schedules represented a monumental methodological paradigm shift. By placing two or more independently programmed manipulanda simultaneously within reach of an organism, choice was redefined from a static, discrete event into a dynamic, continuous allocation of behavior across time. The essential dependent variable was no longer the simple binary count of discrete runs, nor was it the isolated response rate on a single manipulandum. Instead, Herrnstein designated the relative response rate—the proportion of total emitted behavior directed toward a given alternative—as the fundamental behavioral metric. This conceptual breakthrough allowed researchers to systematically map variations in relative environmental inputs directly onto relative behavioral outputs, opening the door to the mathematical formalization of operant preference.
2. Methodological Architecture of the 1961 Seminal Experiment
2.1 Apparatus and Experimental Subjects
Herrnstein’s landmark 1961 experiment utilized adult homing pigeons (Columba livia) as experimental subjects. Pigeons were selected due to their exceptional visual acuity, robust and stereotypic operant response topographies, and long experimental lifespans, which permitted longitudinal research spanning months or years under identical baseline conditions. The experimental apparatus consisted of a custom-fabricated operant conditioning chamber, structurally engineered to eliminate extraneous sensory disruptions. The chamber was housed within a sound-attenuating outer enclosure equipped with a continuous ventilation fan that served the dual purpose of maintaining thermal equilibrium and providing auditory masking against ambient laboratory noise.
The interior operational interface featured two circular translucent plastic response keys mounted horizontally side by side on the front metal panel, separated by a distance of several inches. Each pecking key was backed by an electromechanical microswitch that registered an operant response whenever a pigeon struck the key with a force exceeding a calibrated mechanical threshold (typically around 0.10 to 0.15 Newtons). Behind each translucent key, miniature electrical projection lamps could illuminate the discs with distinct visual stimuli—specifically, different colored lights—allowing the experimenter to establish clear discriminative stimulus conditions for each key. The spatial separation and physical resistance of the keys were precisely calibrated to ensure that an individual response could not inadvertently trigger both switches, and that activating a key required deliberate, focused physical contact.
To establish and maintain motivational stability, the pigeons were subjected to a rigorous nutritional deprivation protocol. Each bird was maintained at approximately 80 percent of its free-feeding body weight throughout the experimental timeline. Birds were weighed daily prior to experimental sessions, with supplemental feeding provided post-session only when necessary to preserve the targeted weight envelope. The primary reinforcing event was the presentation of mixed grain delivered via a solenoid-driven, illuminated grain hopper situated centrally below the two response keys. Whenever a reinforcer was triggered, the key lights extinguished, the chamber house light was darkened, and the grain hopper was elevated and illuminated for a precise temporal duration (typically three seconds), providing the subject with immediate, direct access to food reward while temporarily arresting all active schedule timers.
2.2 Programming Interlocking and Concurrent Reinforcement Schedules
The technical realization of Herrnstein’s experimental design required the engineering of concurrent variable-interval (conc VI VI) reinforcement schedules. In an operant conditioning schedule, a variable-interval schedule dictates that a reinforcer becomes primed and available for delivery only after the passage of a variable, unpredictable interval of time; once this interval elapses, the very next response emitted by the organism on the designated manipulandum immediately collects the reinforcer. In Herrnstein’s concurrent arrangement, two entirely independent VI schedules operated simultaneously—one governing the left response key, and the other governing the right response key. The programming of these temporal intervals was accomplished using classical electromechanical relay circuitry, stepping switches, and continuous perforated paper tape timers running at constant mechanical speeds.
A critical engineering consideration in variable-interval construction was the mathematical distribution of the interval progressions. If intervals within a schedule were arranged arithmetically (e.g., standard linear increments), the organism could detect subtle temporal regularities, leading to cyclic bursts of responding or temporal conditioning similar to fixed-interval scalloping. To prevent such temporal discrimination, intervals were constructed using logarithmic or pseudo-exponential progressions, such as those formalized by Fleshler and Hoffman (1962). Under these progressions, the probability that a reinforcer would become primed in any given second remained essentially constant, regardless of the time that had elapsed since the last reward delivery. This flat hazard function generated remarkably stable, uniform response rates devoid of predictable temporal acceleration or deceleration.
Crucially, the timing mechanisms for the two keys were programmed to operate completely independently of one another. The mechanical timer running the schedule for Key 1 continued to advance and prime reinforcers irrespective of whether the pigeon was pecking Key 1, pecking Key 2, or pausing entirely. If a reinforcer became primed on Key 1 while the bird was actively engaging with Key 2, that reinforcer remained latched in an electromechanical storage relay, waiting indefinitely until the pigeon eventually switched over and executed a single peck on Key 1. This operational independence meant that neglecting either key resulted in an uncollected reward sitting idle on that key, an architectural feature that fundamentally prevented the total behavioral extinction of responding on either alternative.
2.3 Parametric Variations and Experimental Design
Herrnstein’s 1961 study utilized an elegant parametric within-subject design across three individual pigeons, designated as subjects 05, 55, and 231. The primary independent variable was the relative frequency of reinforcement delivered across the two keys. Importantly, Herrnstein engineered these variations while maintaining the total overall reinforcement rate delivered by the two keys combined at a constant aggregate level of 40 reinforcers per hour. This meant that the sum of the reinforcement rates on Key 1 ($R_1$) and Key 2 ($R_2$) was fixed ($R_1 + R_2 = 40\text{ reinforcers/hr}$), while the proportional distribution across the two keys was systematically shifted across different experimental phases.
Over months of continuous testing, Herrnstein subjected the birds to a wide spectrum of relative reinforcement distributions. Conditions ranged from extreme asymmetries—such as concurrent VI 2.25-min VI 45-min (where Key 1 delivered approximately 27 reinforcers per hour and Key 2 delivered only 1.33 reinforcers per hour)—to moderate distributions such as concurrent VI 3-min VI 6-min, and fully symmetrical arrangements like concurrent VI 3-min VI 3-min (20 reinforcers per hour on each key). By sampling across this broad parametric continuum, Herrnstein could evaluate whether behavioral allocation conformed to a continuous mathematical function rather than an idiosyncratic, step-like preference shift.
To establish that the observed response patterns were genuine steady-state functional relationships rather than transient artifacts of learning or behavioral carryover, Herrnstein implemented rigorous steady-state criteria alongside classic ABA reversal phases. Each experimental condition was maintained for dozens of consecutive daily sessions (often 30 to 40 sessions per condition) until visual inspection and statistical criteria confirmed that response rates had achieved asymptotic, stationary equilibrium. When a specific reinforcement ratio was re-tested later in the experiment following intervening conditions, the birds’ behavioral distributions reliably returned to their original values, providing incontrovertible empirical proof of the reversibility and experimental control of the concurrent schedule parameters.
3. The Mechanics of Concurrent Schedules of Reinforcement
3.1 Concurrent Variable-Interval Schedules (Conc VI VI)
Understanding the profound power of concurrent variable-interval schedules requires contrasting them with alternative concurrent arrangements, most notably concurrent fixed-ratio (conc FR FR) schedules. When an organism is confronted with a concurrent FR 50 vs. FR 100 schedule, optimal and empirical behavioral patterns are categorical: the organism will rapidly develop exclusive preference for the smaller ratio (FR 50), allocating 100 percent of its responses to the richer key and zero percent to the leaner key. This occurs because on ratio schedules, every response brings the animal closer to reinforcement; responding on the higher-ratio key simply requires twice the physiological work for identical reward. Ratio schedules yield an “all-or-none” competitive dynamic where the less favorable alternative undergoes complete behavioral extinction.
In stark contrast, concurrent variable-interval schedules actively preserve mixed choice distributions through the mechanics of reinforcer persistence and primed availability. On a VI schedule, responses emitted during the inter-reinforcer interval do not accelerate the arrival of the food; they merely serve to detect when the scheduled temporal interval has elapsed. If a pigeon on a conc VI 1-min VI 3-min schedule spends several consecutive minutes exclusively pecking the richer VI 1-min key, the independent timer on the VI 3-min key will inevitably elapse, latching an uncollected grain reward in the background. At that point, the next single peck directed toward the VI 3-min key will yield an instantaneous reinforcer with an immediate probability of 1.0.
This primed reinforcer mechanism creates a self-correcting negative feedback loop that fundamentally penalizes exclusive preference. The longer an organism stays away from a leaner interval schedule, the higher the momentary probability that a reward is primed and waiting on that neglected key. As soon as the animal samples the neglected manipulandum, it is reliably reinforced on the very first or second response. Consequently, concurrent VI VI schedules mathematically disallow exclusive preference, compelling the organism to allocate its behavioral repertoire between both options to capture all available environmental rewards. It is precisely this structural characteristic that transforms conc VI VI into the premier laboratory vehicle for studying graded preference and behavioral distribution.
3.2 The Changeover Delay (COD) as a Vital Methodological Control
While concurrent VI VI schedules naturally foster distributed responding, early investigations encountered a severe behavioral pathology when the two keys were programmed without transition constraints: the emergence of superstitious switching behavior. When an animal pecks Key 1, pauses, transitions to Key 2, and immediately receives a food reinforcer, the act of switching itself becomes intimately paired with reward delivery. Because operant conditioning is exquisitely sensitive to contiguous temporal relationships, the entire sequence of “peck Key 1, move head rapidly to the right, peck Key 2” becomes reinforced as an integrated, compound behavioral chain. Rather than demonstrating genuine discriminative choice between the two environmental schedules, the pigeon develops a stereotypic alternating motor pattern, indiscriminately shuttling back and forth across the keys at high frequencies.
To eliminate this adventitious reinforcement of switching, Herrnstein and his contemporaries utilized an indispensable methodological device originally formulated by Skinner and polished by Charles Catania: the Changeover Delay (COD). The COD is an interlocking temporal contingency specifying that whenever an organism switches from one response manipulandum to the other, no reinforcer can be delivered on the newly chosen key until a predetermined interval of time has elapsed—typically between 1.5 and 2.0 seconds. If a reinforcer is already primed on Key 2, and the bird pecks Key 1 before switching to Key 2, the bird’s initial peck on Key 2 initiates the 1.5-second COD timer. Any pecks emitted on Key 2 during this 1.5-second buffer are recorded by the apparatus but will not trigger the grain hopper; the primed reinforcer can only be collected by a peck delivered after the COD has fully expired.
The insertion of a COD successfully decouples the motor act of switching from the immediate presentation of reinforcement, cleanly abolishing adventitious behavioral chains. Extensive empirical research has demonstrated that omitting the COD (a COD of 0 seconds) severely degrades behavioral matching, resulting in over-switching and a behavioral distribution that compresses toward an uninformative 50:50 ratio regardless of the underlying reinforcement schedules. Conversely, if the COD is calibrated to an excessively long temporal duration (e.g., 15 to 20 seconds), it transforms the changeover from a mild temporal buffer into an onerous punishment contingency, driving the animal toward overmatching or artificial exclusive preference on whichever key is momentarily richer. Calibrating the COD within an optimal window of 1.5 to 3.0 seconds ensures that the organism’s relative response rates accurately mirror relative reinforcement contingencies.
3.3 Behavioral Dynamics of Switching Between Alternatives
The introduction of the Changeover Delay revealed rich micro-level dynamics governing how animals physically traverse the space between concurrent options. In a concurrent schedule, total behavioral output is not merely a monolithic collection of pecks; it is fundamentally partitioned into discrete “visits” or “stays” on an alternative, punctuated by “changeovers” to the alternate manipulandum. Microstructural analyses show that changeover response rates—the frequency with which an organism abandons its current key to visit the other—are systematic functions of the relative and absolute reinforcement densities scheduled across the environment.
When the two schedules are approximately symmetrical (e.g., VI 2-min vs. VI 2-min), an animal’s stay duration on each key remains brief and roughly equivalent, resulting in frequent, balanced changeovers. However, as the schedule asymmetry increases (e.g., VI 1-min vs. VI 9-min), the organism’s stay duration on the richer alternative lengthens substantially, while visits to the leaner key are brief, highly focused sampling excursions. The animal transitions to the lean key, emits the minimal number of pecks necessary to satisfy the COD and claim any waiting primed reinforcer, and immediately returns to the richer patch. The energetic and temporal overhead associated with the physical transit—the changeover cost—acts as a localized friction that modulates the temporal threshold required for an animal to abandon its current foraging baseline.
Furthermore, microstructural analysis reveals notable post-reinforcement pausing patterns within concurrent setups. Immediately following the consumption of a reinforcer from the grain hopper, the probability of an immediate changeover to the opposite key spikes dramatically. This phenomenon reflects an adaptive foraging logic: because the animal has just exhausted the reinforcer on the currently chosen key, the momentary probability of obtaining an instantaneous second reinforcer on that identical key drops toward zero (governed by the lower bound of the interval distribution). Concurrently, the unvisited key has had uninterrupted time to accumulate a primed reward. The post-reinforcement pause thus functions as a strategic pivot point where the animal re-evaluates the momentary values of the two alternatives, driving organized, non-random transition sequences.
4. Mathematical Formulation and Derivation of the Simple Matching Law
4.1 The Proportional Matching Equation
Herrnstein’s analysis of his 1961 empirical data culminated in a remarkably concise and elegant algebraic formulation. When he plotted the relative frequency of responding on a chosen key against the relative frequency of reinforcement obtained on that same key, the data points aligned directly along an identity line. Herrnstein formalized this observation as the Simple Matching Law, mathematically defined by the equation:
$$\frac{B_1}{B_1 + B_2} = \frac{R_1}{R_1 + R_2}$$
In this formulation, $B_1$ and $B_2$ denote the absolute behavioral outputs (measured as total responses emitted, or pecks) on Manipulandum 1 and Manipulandum 2, respectively. The terms $R_1$ and $R_2$ represent the absolute obtained reinforcement frequencies (measured as reinforcers delivered per unit time) harvested from Manipulandum 1 and Manipulandum 2. The quotient $B_1 / (B_1 + B_2)$ represents the relative response rate directed toward Option 1, bounded strictly between 0.0 and 1.0. Similarly, the quotient $R_1 / (R_1 + R_2)$ represents the relative reinforcement rate delivered by Option 1, also normalized on the unit interval [0, 1].
The philosophical and mathematical significance of this equation lies in its striking parsimony. Unlike most psychological models of the era, Herrnstein’s simple matching equation possessed no free empirical parameters, no scaling constants, and no fitted subjective weighting exponents. It posited that operant behavioral allocation is directly proportional to environmental reinforcement allocation. When graphed on a standard Cartesian coordinate system with relative reinforcement on the abscissa and relative responding on the ordinate, the matching law predicts that empirical data will fall cleanly along a 45-degree diagonal line extending from the origin $(0,0)$ to the opposite apex $(1,1)$. If an organism secures 70 percent of its total reinforcement from Key 1, it will allocate exactly 70 percent of its total behavioral responses to Key 1.
4.2 The Ratio Formulation of the Matching Law
While the proportional matching equation provides an intuitive visualization of relative choice bounded between zero and one, it possesses mathematical limitations, notably that proportional data suffer from structural collinearity—because the relative values must sum to 1.0, the variance of $B_1 / (B_1 + B_2)$ is mechanically constrained as it approaches either asymptote. To bypass these limitations and facilitate deeper algebraic analysis, researchers reformulated the matching law into an equivalent ratio formulation. Through straightforward algebraic manipulation of the proportional equation, one can derive:
$$\frac{B_1}{B_2} = \frac{R_1}{R_2}$$
The ratio formulation isolates relative preference directly: the ratio of responses allocated between the two alternatives is an identical match to the ratio of reinforcers obtained from those alternatives. This format offers immense analytical advantages because it decouples preference from total behavioral output, allowing researchers to study choice across environments characterized by radically divergent overall response rates. Furthermore, the ratio formulation can be readily converted into a linear function via a logarithmic transformation:
$$\log\left(\frac{B_1}{B_2}\right) = \log\left(\frac{R_1}{R_2}\right)$$
In this logarithmic space, the matching relationship is expressed as a simple linear equation of the standard form $y = mx + c$, where the slope $m = 1.0$ and the intercept $c = 0.0$. This logarithmic expression provided the empirical foundation for modern linear regression diagnostics in behavioral psychology. However, working within the ratio framework introduces a specific mathematical challenge: if an organism emits zero responses on an alternative ($B_2 = 0$), or if a schedule yields zero reinforcers ($R_2 = 0$), the ratio involves division by zero, resulting in undefined mathematical singularities. Consequently, ratio models require organisms to sample both manipulanda, an outcome inherently fostered by concurrent variable-interval schedules.
4.3 Herrnstein’s Hyperbolic Absolute Rate Equation
While the proportional matching law successfully accounted for choice between two discrete keys, Herrnstein recognized that a universal law of behavior must also account for performance under single-operant schedules, where an animal responds on only one manipulandum. In a single-operant environment (e.g., a simple VI schedule), how does an organism determine its absolute response rate without a second visible key? Herrnstein proposed a brilliant conceptual synthesis: *an organism is never choosing in a vacuum*. Even in an experimental chamber equipped with only one mechanical lever, the animal is continuously deciding between engaging with that lever versus allocating behavior to unmeasured, extraneous background activities—such as grooming, exploring, sniffing the air, or resting.
To capture this reality, Herrnstein expanded the matching principle into his famous Hyperbolic Absolute Rate Equation for single schedules of reinforcement:
$$B_1 = \frac{k \cdot R_1}{R_1 + R_e}$$
In this equation, $B_1$ is the absolute response rate on the single programmed operant manipulandum, and $R_1$ is the rate of reinforcement delivered by that schedule. The parameter $k$ represents the organism’s total behavioral capacity—the theoretical maximum asymptotic response rate that the animal’s motor system could physically emit if it dedicated 100 percent of its active behavioral budget to that single response. The parameter $R_e$ denotes the aggregate rate of extraneous reinforcement—the total volume of unprogrammed, background reinforcing events intrinsically provided by the environment (e.g., visual novelty, motor relief, self-grooming, sensory feedback).
Mathematically, the hyperbolic equation is a rectangular hyperbola that passes through the origin. When the programmed reinforcement rate $R_1$ is low, small increases in scheduled reward produce massive escalations in behavioral output. However, as $R_1$ grows large relative to the extraneous reinforcement $R_e$, the term $R_1 / (R_1 + R_e)$ asymptotically approaches 1.0, causing the response rate $B_1$ to plateau smoothly toward the motor capacity limit $k$. Herrnstein’s hyperbolic equation demonstrated that the matching law was not merely a narrow description of concurrent keys, but a profound, universal framework governing all operant behavior: absolute responding on any single schedule is simply a matching response against the broader context of extraneous environmental reinforcement.
5. Empirical Findings and the Robustness of the Matching Phenomenon
5.1 Analysis of Herrnstein’s Original 1961 Empirical Data
Herrnstein’s 1961 experimental results provided a demonstration of quantitative precision that electrified the behavioral research community. Plotting the relative response rates of his three pigeon subjects across dozens of experimental sessions yielded empirical data points that clustered tightly around the theoretical 45-degree identity line. When submitted to quantitative correlation analysis, the data yielded Pearson correlation coefficients ($r$) exceeding 0.98 for all three subjects. The empirical fit was remarkably robust, showing that variation in the programmed relative reinforcement rate accounted for virtually all the variance observed in the pigeons’ relative response allocation.
What made these findings particularly compelling was their stability across radically different temporal schedules. Herrnstein varied not only the relative ratio between Key 1 and Key 2, but demonstrated that the matching relationship held constant whether the overall schedule density was fast or slow, provided that relative contingencies were preserved. Modern re-analyses of Herrnstein’s original tabular data using contemporary regression metrics continue to validate his conclusions. While minor idiosyncratic deviations were visible in individual session baselines, the aggregated steady-state data demonstrated that an animal’s nervous system was capable of tracking, computing, and matching relative environmental reward rates with an accuracy rivaling sensory psychophysics.
Herrnstein’s findings dealt a major blow to classic Hullian and early Skinnerian assertions that an absolute quantity of drive or localized reinforcement histories governed isolated actions. Instead, the 1961 data decisively proved that an operant response rate is fundamentally relativistic. An organism does not possess an immutable response rate for a VI 2-min schedule; pecking on a VI 2-min schedule will be exceptionally high if the concurrent alternative is a sparse VI 30-min schedule, but that identical VI 2-min schedule will yield an anemic trickle of responses if placed concurrently alongside an opulent VI 15-second schedule. Response strength, Herrnstein proved, is inherently relative to the contextual reinforcement matrix.
5.2 Systematic Replications across Diverse Taxa
Following the publication of Herrnstein’s monograph, experimental psychologists embarked on decades of systematic cross-taxa replications to determine whether the matching law was an avian idiosyncrasy or a fundamental biological invariant. In rodent laboratories, researchers deployed dual-lever and dual-nose-poke operant conditioning chambers to test laboratory rats (Rattus norvegicus). Utilizing liquid sucrose, food pellets, or intracranial self-stimulation as rewards, rodent studies overwhelmingly corroborated Herrnstein’s findings, showing that mammalian motor topographies adhered closely to matching principles under appropriate concurrent interval schedules and changeover delays.
Subsequent extensions ventured into primate research, testing rhesus macaques, baboons, and chimpanzees under sophisticated multi-operant protocols. Primates demonstrated matching across varied manual tasks, joystick tracking, and touchscreen visual displays. Concurrently, human behavioral researchers demonstrated matching in controlled laboratory choice environments, showing that human adults and children distribute button-presses, token-earning responses, and visual fixations in precise proportion to relative reinforcement schedules, provided that verbal self-instructions or explicit rule-governed behavior do not artificially override direct schedule contact.
The taxonomic breadth of matching expanded further as behavioral ecologists and comparative psychologists tested diverse species, including dairy cows, domestic dogs, corvids, and teleost fish. Beyond the laboratory chamber, behavioral ecologists observed that wild foraging animals operating in natural ecosystems allocate their foraging time across discrete resource patches in direct proportion to the relative caloric yields of those patches. Whether observing a flock of mallard ducks distributing themselves across two continuous bread-tossing feeding stations on an open lake, or bees allocating visits to competing flower patches differing in nectar concentration, the matching law emerged as a universal biological principle governing behavioral allocation across phylogenetic history.
5.3 Matching Across Different Reinforcer Dimensions
While Herrnstein’s initial 1961 experiment focused exclusively on variations in the rate (frequency) of reinforcement, subsequent investigators sought to determine whether the matching principle generalized to other fundamental dimensions of reward. Animals in natural ecologies make choices not only between patches that yield food more or less frequently, but between patches offering rewards that vary in physical size, caloric density, temporal delay, and qualitative palatability. A truly universal model of choice required accommodating these multi-dimensional attributes into a single quantitative framework.
Experimental arrangements systematically manipulated reinforcer magnitude (e.g., varying grain hopper duration from 1.5 to 6.0 seconds, or delivering different volumes of sucrose solution) while holding reinforcement rates equal. These studies revealed that relative response rates matched relative reinforcer magnitudes ($M_1 / M_2$):
$$\frac{B_1}{B_2} = \frac{M_1}{M_2}$$
Similarly, when researchers varied the immediacy of reinforcement—the inverse of the temporal delay ($D$) intervening between a response and reward delivery—behavioral allocation matched the relative immediacy of reinforcement ($1/D_1$ vs. $1/D_2$), demonstrating that delayed reinforcers exert diminished control over behavioral output in precise quantitative proportions.
To integrate these divergent dimensions into a unified mathematical structure, William Baum and other quantitative theorists formulated the Multidimensional Matching Equation. In this expanded model, the absolute value ($V$) of an alternative is conceptualized as the product of its individual physical dimensions: rate ($R$), magnitude ($M$), immediacy ($1/D$), and qualitative preference ($Q$). The overarching matching relationship is thus expressed as:
$$\frac{B_1}{B_2} = \left(\frac{R_1}{R_2}\right) \times \left(\frac{M_1}{M_2}\right) \times \left(\frac{D_2}{D_1}\right) \times \left(\frac{Q_1}{Q_2}\right)$$
This multidimensional synthesis demonstrated that the matching law is essentially an equation of economic exchange value: an organism balances its behavioral expenditure against the aggregated subjective utility produced by compounding rate, size, speed, and quality of reward.
6. Systematic Deviations: Undermatching, Overmatching, and Bias
6.1 Undermatching: Prevalence and Underlying Causes
As laboratories across the world conducted hundreds of concurrent schedule experiments, it became increasingly apparent that while empirical data closely mirrored Herrnstein’s equation, systematic quantitative deviations were remarkably common. The most pervasive of these deviations is undermatching. Undermatching occurs when an organism’s behavioral allocation is *less extreme* than the allocation of reinforcement delivered by the environment. In an undermatching scenario, the animal responds to the richer alternative less than predicted by the matching law, and responds to the leaner alternative more than predicted, effectively compressing the behavioral ratio toward indifference (a 50:50 distribution).
The primary and most thoroughly documented cause of undermatching is an inadequate Changeover Delay (COD). If the COD is calibrated to an insufficient temporal interval (e.g., 0.5 seconds or less), or omitted entirely, accidental adventitious reinforcement of switching chains contaminates the choice baseline. When this occurs, the pigeon or rat continues to visit the lean key far more frequently than warranted by that key’s scheduled rate of reward, because the transition itself carries conditioned reinforcement value. Even in studies utilizing standard CODs, minor sensory discrimination failures can generate undermatching: if the visual stimuli marking Key 1 and Key 2 are perceptually similar, or if an animal experiences momentary lapses in stimulus control, behavioral distributions inevitably drift toward the middle.
Another systemic driver of undermatching is behavioral inertia and imperfect sampling dynamics. Organisms do not possess omniscient knowledge of scheduled reinforcement rates; they must continually harvest information through active sampling. Because the environmental landscape in nature is non-stationary, evolutionary selection favors animals that retain a residual tendency to intermittently explore and sample poorer alternatives, ensuring they do not miss sudden, profitable shifts in resource availability. In a logarithmic plot of choice versus reinforcement, undermatching manifests visually as a fitted regression slope that is systematically flatter than the theoretical identity slope of 1.0 (typically ranging between 0.75 and 0.90).
6.2 Overmatching: Conditions and Theoretical Significance
The inverse deviation from standard matching is overmatching. Overmatching is characterized by behavioral allocation that is *more extreme* than the scheduled allocation of reinforcement. When overmatching occurs, the organism over-allocates its behavior to the richer alternative beyond what proportional matching predicts, while systematically under-allocating its responses to the leaner alternative. On a logarithmic choice plot, overmatching appears as a regression slope strictly greater than 1.0.
Overmatching is observed almost exclusively under conditions where the physical, temporal, or energetic costs of switching between manipulanda are exceptionally high. In standard operant chambers, the two keys are separated by only a few inches, allowing the animal to switch between them with negligible physical effort. However, if researchers alter the experimental architecture by introducing substantial physical barriers—such as placing the keys at opposite ends of a large chamber, requiring the animal to navigate a complex hurdle, or inserting an extended Changeover Delay of 10 to 20 seconds—overmatching reliably emerges.
From an ecological and evolutionary perspective, overmatching represents a highly adaptive behavioral optimization under high travel costs. In natural foraging, traversing an extensive, predator-exposed distance to sample a distant, impoverished patch is energetically counterproductive. When travel costs are severe, an organism maximizes its net energetic gain by remaining entrenched within the known richer patch, abandoning it only rarely. The emergence of overmatching demonstrates that an organism’s behavioral sensitivity to reinforcement ratios is fundamentally bounded and calibrated by the ecological architecture of the surrounding environment.
6.3 Systematic Bias: Asymmetries in Choice Allocation
The third major empirical deviation from the simple matching law is systematic bias. Unlike undermatching and overmatching, which describe deviations in the *slope* (sensitivity) of the behavioral response to reinforcement, bias refers to a constant, schedule-independent preference for one alternative over another. When bias is present, an animal consistently allocates a higher proportion of its behavior toward a specific manipulandum, regardless of how the relative reinforcement rates are configured across experimental phases. On a logarithmic choice plot, bias manifests as a vertical displacement—an intercept shift away from the origin $(0,0)$.
Bias can arise from a multitude of trivial physical factors or profound qualitative differences. Common methodological causes include motoric asymmetries inherent to the organism (equivalent to handedness or visual field dominance), microswitch resistance discrepancies (where one key requires slightly less physical force to actuate than the other), or subtle differences in the illumination intensity of the response keys. If Key 1 is physically easier to depress than Key 2, the bird will execute more pecks on Key 1 across all conditions, producing a constant upward intercept shift.
More importantly, bias serves as an extraordinarily sensitive quantitative assay for evaluating non-rate reward preferences. If Key 1 delivers high-protein hemp seed while Key 2 delivers standard grain, the pigeon will exhibit a massive, consistent bias toward Key 1. By holding physical manipulanda parameters constant and introducing qualitative variations in reinforcer flavor, temperature, or sensory modality, researchers can utilize the magnitude of the bias intercept to quantify the exact subjective exchange value of diverse biological commodities that cannot be compared along a simple physical dimension like rate or mass.
7. The Generalized Matching Law: Formulation and Interpretation
7.1 Baum’s Power Function Formulation
Recognizing that undermatching, overmatching, and systematic bias were ubiquitous, lawful features of choice data rather than random experimental noise, William M. Baum published a landmark 1974 theoretical paper introducing the Generalized Matching Law (GML). Baum proposed modeling choice as a power function, incorporating two free empirical parameters to capture sensitivity and bias directly. Baum’s power function equation is formulated as:
$$\frac{B_1}{B_2} = b \left(\frac{R_1}{R_2}\right)^a$$
In this equation, the behavioral ratio $(B_1 / B_2)$ remains the dependent choice variable, and the obtained reinforcement ratio $(R_1 / R_2)$ remains the primary independent variable. However, the relationship is now moderated by two empirical parameters: $a$ and $b$. The exponent $a$ represents sensitivity to reinforcement ratios, while the multiplier $b$ represents constant bias toward Alternative 1.
The true clinical and analytical brilliance of Baum’s generalized matching law becomes evident when the equation is subjected to a logarithmic transformation. Taking the logarithm of both sides of the power function yields a straightforward linear relationship:
$$\log\left(\frac{B_1}{B_2}\right) = a \cdot \log\left(\frac{R_1}{R_2}\right) + \log(b)$$
This log-linear transformation directly matches the standard equation for a straight line ($y = mx + c$). When behavioral data are plotted in logarithmic ratio coordinates, the slope of the best-fitting linear regression line is equal to the sensitivity parameter $a$, and the y-intercept of the line is equal to $log(b)$. If $a = 1.0$ and $b = 1.0$ (meaning $log(b) = 0$), Baum’s generalized matching law reduces precisely to Herrnstein’s original 1961 simple matching law. The GML does not discard Herrnstein’s original insight; rather, it embeds it within a sophisticated psychophysical framework capable of accommodating the full spectrum of empirical choice behavior.
7.2 Statistical Parameter Estimation Techniques
The log-linear formulation of the Generalized Matching Law allowed behavioral analysts to apply standard Ordinary Least Squares (OLS) linear regression to quantify behavioral datasets. Researchers collect steady-state response counts and obtained reinforcers across multiple schedule conditions for an individual subject, calculate the logarithmic ratios $\log(B_1 / B_2)$ and $\log(R_1 / R_2)$ for each condition, and fit a regression line through the coordinate points. The resulting slope estimates $a$, the intercept estimates $log(b)$, and the coefficient of determination ($R^2$) reveals the proportion of behavioral variance explained by the schedule contingencies.
However, applying simple OLS regression to operant datasets can introduce statistical complexities. Because operant conditioning experiments typically collect repeated-measures data from a small cohort of subjects exposed to longitudinal within-subject designs, the assumption of independent and identically distributed errors can be compromised. Individual subjects often display idiosyncratic baseline variances, introducing potential autocorrelation across longitudinal phases. Furthermore, ratio coordinates are susceptible to heteroscedasticity, where the variance of error terms fluctuates depending on whether the reinforcement ratio is balanced or pushed to extreme margins.
To overcome these regression challenges, modern quantitative behavior analysis increasingly employs nonlinear mixed-effects models and hierarchical Bayesian modeling. These modern statistical techniques permit the simultaneous estimation of fixed effects (representing the species-level or population-level sensitivity $a$ and bias $b$) and random effects (capturing the unique parametric deviations of individual subjects). By modeling the hierarchical structure of the data directly, researchers obtain robust, unbiased parameter estimates that accommodate unbalanced datasets, repeated measures, and minor session-to-session fluctuations without sacrificing statistical power.
7.3 Theoretical Implications of the Sensitivity Parameter
The sensitivity parameter $a$ in the Generalized Matching Law is not merely an arbitrary curve-fitting scalar; it serves as a profound psychophysical metric indexing an organism’s behavioral discrimination of reinforcement contingencies. In classical sensory psychophysics, Weber’s Law and Fechner’s formulations demonstrated that the perceived intensity of a physical stimulus scales logarithmically with its physical magnitude. The sensitivity parameter $a$ functions as an operant analogue to the psychophysical Weber fraction, reflecting how accurately an organism’s nervous system registers and translates variations in environmental reward frequencies into motor commands.
When $a = 1.0$, the organism exhibits perfect sensitivity, demonstrating optimal, proportional tracking of the reinforcement environment. When $a < 1.0$ (undermatching), it indicates a perceptual or behavioral deficit in sensitivity: the organism is under-responsive to shifts in relative reward contingencies, behaving as though the two alternatives are more similar than they objectively are. Sensitivity values are profoundly degraded by cognitive load, environmental distractors, ambiguous discriminative stimuli, neurological insults, or pharmacological manipulations (such as dopamine receptor antagonists). When an environment becomes excessively noisy or cognitively complex, sensitivity $a$ systematically deteriorates toward zero, where behavioral allocation becomes completely unresponsive to environmental contingencies.
Conversely, parameter $a$ can be harnessed clinically and diagnostically as a sensitive behavioral biomarker. Clinical studies evaluating individuals with Autism Spectrum Disorder (ASD), Attention-Deficit/Hyperactivity Disorder (ADHD), or traumatic brain injuries have demonstrated that these populations frequently exhibit quantifiable alterations in sensitivity parameters on concurrent schedule tasks. By assessing parameter $a$, translational clinicians can objectively measure an individual’s behavioral plasticity and responsiveness to reinforcement modifications, providing an empirical baseline for evaluating cognitive therapies and pharmacological interventions.
8. Theoretical Mechanisms: Maximization versus Melioration
8.1 Molecular Maximization Theories
The discovery of the matching law ignited a fiery, decades-long theoretical controversy regarding the underlying behavioral mechanisms that generate matching. The central debate focused on whether matching is a primary, fundamental biological law, or merely an emergent mathematical byproduct of an organism striving to *maximize* its reward intake. The earliest mechanistic alternatives were Molecular Maximization Theories, championed prominently by Charles Shimp and colleagues during the late 1960s and 1970s.
Molecular maximization posits that organisms operate strictly on micro-level temporal contingencies, continuously selecting whichever individual response alternative possesses the highest *momentary probability of reinforcement* at the precise microsecond the action is emitted. Because concurrent variable-interval schedules allow uncollected reinforcers to accumulate over time, an animal that has spent several seconds pecking Key 1 faces an environment where the momentary probability of Key 1 delivering a reward is slowly decaying or stationary, while the probability of Key 2 harboring an uncollected reward is compounding with each passing second. Eventually, the momentary probability on Key 2 surpasses that of Key 1, triggering a switch.
Proponents of molecular maximization developed intricate mathematical models demonstrating that if an organism executes momentary probability calculations and switches responses precisely when the local likelihood of reward favors the competing key, the macroscopic, aggregated sum of these micro-decisions will closely approximate matching law distributions. However, subsequent empirical research revealed fatal contradictions in pure molecular maximization. Experimenters engineered specialized, non-linear concurrent schedules—such as synthetic probability schedules and inter-reinforcer interval clamps—where molecular maximization explicitly dictated behavioral patterns that diverged sharply from matching. Under these specialized contingencies, animals consistently adhered to macroscopic matching, willfully abandoning the molecularly optimal response path.
8.2 Molar Maximization Theories
In direct opposition to molecular models, Howard Rachlin, Leonard Green, and other behavioral economists formulated Molar Maximization Theories. Grounded heavily in neoclassical economics, molar maximization asserts that organisms do not calculate microscopic, momentary probabilities; instead, they evaluate global, aggregate outcomes over extended temporal horizons. According to this framework, an organism’s behavioral allocation across hours or days is designed to maximize the total, absolute volume of reinforcement harvested, subject to the organism’s total behavioral budget and the physical constraints of the environment.
Molar theorists argued that matching was simply an accidental byproduct of an animal optimizing its molar utility under standard concurrent variable-interval schedules. Because concurrent VI VI schedules utilize independent timers, the schedule architecture inherently ensures that an animal can harvest nearly 100 percent of scheduled reinforcers only if it distributes its responses across both keys in a balanced manner. If the animal were to allocate 100 percent of its behavior to Key 1, it would forfeit all reinforcers sitting on Key 2; conversely, allocating behavior proportionally captures the vast majority of rewards on both keys. Thus, molar maximization theorists maintained that the matching law was nothing more than an organism’s rational economic optimization strategy.
However, the molar maximization hypothesis was ultimately undermined by decisive empirical counter-demonstrations. Researchers engineered concurrent schedules pairing a variable-interval schedule with a variable-ratio schedule (conc VI VR), or complex feedback schedules where matching behavior directly reduced the overall aggregate reinforcement rate. Under these paradigms, molar maximization dictated that the organism should allocate the overwhelming bulk of its behavior to the ratio schedule to maximize total reward intake. Astoundingly, experimental animals persistently adhered to matching distributions, sacrificing substantial volumes of available food reward to maintain matching equilibrium. These demonstrations conclusively proved that organisms are not molar optimizers; matching was driving behavior, even when it led to profound, global economic sub-optimality.
8.3 Herrnstein and Prelec’s Theory of Melioration
Confronted with the structural failures of both molecular and molar maximization, Richard Herrnstein, in close collaboration with economist and mathematician Drazen Prelec, formulated the Theory of Melioration. Derived from the Latin root *meliorare* (meaning “to make better”), melioration posited that choice is driven neither by momentary molecular probabilities nor by global molar optimization, but by an ongoing, dynamic process of localized comparison. Organisms do not maximize; they meliorate.
The mathematical engine of melioration centers on the concept of the local reinforcement rate. The local reinforcement rate on an alternative is defined as the number of reinforcers harvested from that alternative divided by the actual behavioral time or energy expended directly on that alternative ($R_i / B_i$), excluding time spent on other options. The melioration hypothesis asserts a fundamental, simple behavioral rule: an organism will continuously shift its behavioral allocation toward whichever alternative currently provides the higher local rate of reinforcement. If Alternative 1 yields 10 reinforcers per minute of time spent pecking it, while Alternative 2 yields only 5 reinforcers per minute of time spent on it, the organism will allocate more time and behavior to Alternative 1.
As the organism re-allocates more behavior toward Alternative 1, its increased behavioral investment naturally drives down the local reinforcement rate on that key, due to the diminishing returns inherent in variable-interval schedules. Concurrently, reducing behavior on Alternative 2 causes its local reinforcement rate to rise as uncollected rewards accumulate. This dynamic process of continuous shifting inevitably terminates at a stationary behavioral equilibrium where the local reinforcement rates across all available alternatives become precisely equalized:
$$\frac{R_1}{B_1} = \frac{R_2}{B_2}$$
Through simple algebraic cross-multiplication, this equalization of local rates rearranges into:
$$\frac{B_1}{B_2} = \frac{R_1}{R_2}$$
Melioration provides the ultimate mechanistic explanation: the Matching Law is the inevitable mathematical equilibrium of local rate equalization. Crucially, Herrnstein and Prelec demonstrated that because melioration operates strictly on local rate differentials, it operates blindly with respect to global, overall outcomes. In environments where allocating behavior to a locally attractive option degrades the global payoffs of both options (a dynamic known as a “primrose path”), melioration inexorably drives the organism into severe sub-optimality, addiction traps, and economic ruin—precisely matching real-world behavioral pathologies observed across species.
9. Extensions of the Matching Law: Time Allocation and Concurrent Chains
9.1 Time Allocation Matching
While Herrnstein’s original 1961 formulation measured behavior through discrete response tallies ($B_1$ and $B_2$), William Baum recognized that in many real-world and laboratory contexts, behavior is continuous rather than discrete. Foraging animals do not merely tally discrete peck or lever-press counts; they spend continuous durations of time residing within, navigating, and inspecting different ecological environments. In 1974, Baum proposed that the matching law could be formulated more universally by substituting discrete response frequencies with measures of continuous Time Allocation ($T$):
$$\frac{T_1}{T_1 + T_2} = \frac{R_1}{R_1 + R_2}$$
In this temporal formulation, $T_1$ and $T_2$ denote the total duration of time an organism commits to residing within or actively attending to Alternative 1 and Alternative 2 during an experimental session. Baum demonstrated that time allocation matching exhibits the same profound quantitative invariants as discrete response matching, frequently yielding even cleaner empirical fits with less undermatching. Because allocating time is the universal behavioral denominator—all organisms possess an inescapable, non-negotiable budget of 24 hours per day—time allocation provided a universal currency capable of bridging disparate operant topographies.
Time allocation matching proved immensely powerful in evaluating behaviors that do not possess neat, discrete motor units. Complex tasks such as sustained vigilance monitoring, parental care, social grooming, predatory stalking, and human multi-tasking cannot be parsed into identical pecks; however, they can be flawlessly quantified through continuous time tracking. Subsequent comparative research demonstrated that when organisms navigate continuous environments, relative time investment matches relative obtained reinforcement with extraordinary precision, affirming that temporal expenditure is the primary medium through which biological valuation is manifested.
9.2 Concurrent-Chain Schedules and Conditioned Reinforcement
To investigate how organisms make commitments and evaluate delayed, probabilistic, or conditioned rewards, behavioral analysts developed the Concurrent-Chain Schedule apparatus, pioneered by David Autor and refined extensively by Edmund Fantino. A concurrent-chain schedule fractures the choice paradigm into two distinct, sequential temporal phases: an initial link (the choice phase) and a terminal link (the outcome phase). In the initial link, two concurrent keys are illuminated (e.g., both white), operating under independent concurrent VI VI schedules. Pecks on these initial-link keys do not deliver primary food reinforcement; instead, satisfying the variable-interval requirement on one of the keys immediately triggers a transition: both keys change color, and the organism enters the designated terminal link.
Once an organism enters a terminal link, the alternative option is completely turned off, locking the animal into its chosen path. The animal must then satisfy the specific schedule programmed within that terminal link (e.g., a fixed delay, a variable interval, or an aversive stimulus) before finally receiving primary reinforcement. Because responding during the initial link is maintained entirely by the opportunity to enter a specific terminal link, the relative response rate during the initial link serves as an uncontaminated, direct quantitative assay of the conditioned reinforcing strength of the stimulus associated with that terminal link.
To account for choice in concurrent chains, Edmund Fantino formulated the Delay-Reduction Hypothesis, which directly extends the matching law to conditioned reinforcement. Fantino demonstrated that the conditioned reinforcing value of an outcome is not determined by its absolute temporal duration, but by the degree to which its onset signals a reduction in the overall time remaining to primary reinforcement, relative to the average baseline delay of the entire experimental context. Concurrent-chain architectures provided the empirical bedrock for modern behavioral investigations into intertemporal choice, temporal discounting, impulsivity, and the mathematical modeling of conditioned value transfer.
9.3 Contextual Matching and Multiple Alternative Choices
While binary choice paradigms dominated the initial decades of research, natural environments rarely present organisms with an isolated dichotomy. Foraging ecologies, modern commercial retail spaces, and digital social media landscapes present agents with an expansive matrix of concurrent alternatives ($N ge 3$). The simple matching law readily scales mathematically to accommodate any finite number of concurrent options through generalized summation:
$$\frac{B_i}{\sum_{j=1}^{N} B_j} = \frac{R_i}{\sum_{j=1}^{N} R_j}$$
This contextual expansion demonstrates that the behavioral allocation dedicated to any specific option $i$ is inherently determined by the total aggregate reinforcement pool generated by all $N$ available options combined. If a third, highly lucrative alternative is injected into an existing two-choice environment, it inevitably cannibalizes behavioral investment from both original options, redistributing total response allocation to maintain proportional equilibrium across the broader contextual array.
Importantly, empirical evaluations of multi-alternative matching have provided critical insights into classical economic axioms, notably the Independence of Irrelevant Alternatives (IIA) and the Luce Choice Axiom. IIA asserts that the relative preference ratio between Option 1 and Option 2 should remain entirely invariant regardless of whether a third alternative (Option 3) is introduced or removed from the choice set. While the simple multi-alternative matching equation assumes strict adherence to IIA, empirical investigations have uncovered subtle context-dependent violations—such as asymmetric dominance effects, phantom decoy alternatives, and similarity clustering—demonstrating that reinforcement value is dynamically modulated by the structural topology of the choice matrix.
10. Behavioral Economics and the Synthesis of Choice
10.1 The Intersection of the Matching Law and Consumer Demand Theory
During the late 1970s and 1980s, behavioral analysts and microeconomists converged to form the interdisciplinary domain of Behavioral Economics, with the matching law serving as a vital conceptual catalyst. In this synthesis, operant conditioning arrangements were mapped directly onto microeconomic consumer demand theory. The operant response requirement (e.g., the number of pecks or lever presses mandated by a schedule) was formally operationalized as price ($P$), the rate of reinforcer delivery was operationalized as consumption ($Q$), and total session responding was operationalized as expenditure.
This economic framing allowed researchers to evaluate matching dynamics through the lens of own-price elasticity and cross-price elasticity of demand. When an animal chooses between two concurrent reinforcers, its behavioral allocation is governed not merely by scheduled reward rates, but by the degree of economic *substitutability* between the commodities. Two reinforcers are defined as perfect economic substitutes if an increase in the price of Reinforcer A causes an immediate compensatory surge in the consumption of Reinforcer B (such as two identical carbohydrate grain pellets). Conversely, if the reinforcers are complementary (such as food and water), an increase in the work price of food causes consumption of both food and water to decline concurrently.
Integrating economic demand curves with the Generalized Matching Law revealed that the sensitivity parameter $a$ is intimately linked to commodity substitutability. When concurrent options offer completely substitutable rewards, an organism exhibits high sensitivity to price differentials, driving parameter $a$ toward or above 1.0. However, when concurrent alternatives represent non-substitutable, functionally independent commodities (e.g., caloric nourishment versus wheel-running access), the organism cannot freely trade off consumption of one for the other; consequently, behavioral allocation becomes highly inelastic, driving sensitivity parameter $a$ toward zero. The matching law was thus transformed from a purely descriptive frequency equation into an expressive model of resource allocation under economic constraints.
10.2 The Concept of Behavioral Surplus and Closed versus Open Economies
A crucial empirical breakthrough in operant behavioral economics was the realization that laboratory choice experiments differ radically depending on whether they are conducted in open economies or closed economies, a distinction comprehensively formalized by George Collier and William Hursh. In an *open economy*, experimental sessions are brief (e.g., 30 to 60 minutes); the animal earns only a fraction of its total daily nutritional intake inside the apparatus and receives supplementary post-session feedings in its home cage regardless of experimental performance. In a *closed economy*, the subject resides inside the operant conditioning chamber 24 hours a day, 7 days a week; all biologically required food and water must be earned strictly through schedule interaction, with zero external feeding subsidies.
The economic architecture of the experimental chamber exerts a dramatic impact on matching law parameters, particularly the extraneous reinforcement term $R_e$ in Herrnstein’s hyperbolic formulation. In an open economy, the guaranteed arrival of free post-session food represents a massive reservoir of unprogrammed, extraneous background reinforcement ($R_e$). Because the animal’s ultimate survival is insulated by home-cage feeding, its demand for within-session reinforcers is highly elastic. If schedule prices rise or asymmetries widen, the animal can simply cease responding, wait for the session to terminate, and consume its post-session rations. This dynamic suppresses within-session behavioral output and frequently induces significant undermatching.
In stark contrast, when subjects operate within a closed economy, the extraneous reinforcement term $R_e$ is driven toward zero because the chamber encompasses the animal’s entire universe. Under closed conditions, demand for life-sustaining reinforcers becomes remarkably inelastic: the animal *must* complete the requisite schedule requirements to survive. Consequently, animals in closed economies exhibit extraordinary persistence, tolerating massive ratio requirements and displaying robust behavioral matching distributions that reflect genuine ecological trade-offs, fully aligning operant choice with classical budgetary indifference curves.
10.3 Temporal Discounting, Impulsivity, and Hyperbolic Valuation
Perhaps the most widespread and culturally impactful extension of Herrnstein’s work occurred through its application to intertemporal choice and temporal discounting. In everyday decision-making, agents must constantly choose between smaller, immediate rewards (e.g., recreational leisure, smoking, consumer spending) and larger, delayed rewards (e.g., academic attainment, physical health, retirement investments). Standard neoclassical economics historically modeled these choices using exponential discount functions derived from compound interest formulas, assuming that rational agents discount future rewards at a constant percentage per unit of time.
However, real human and animal subjects display dramatic, predictable preference reversals: an individual may rationally prefer $105 delivered in 365 days over$100 delivered in 364 days, yet when the choice arrives in the present moment, reverse their preference to choose an immediate $100 today over$105 tomorrow. Drawing directly from Herrnstein’s matching principles and the work of George Ainslie, researchers demonstrated that temporal discounting does not follow an exponential curve; instead, it is governed by a Hyperbolic Discount Function, formally articulated by the Ainslie-Herrnstein model:
$$V = \frac{A}{1 + k \cdot D}$$
In this classic equation, $V$ represents the subjective present value of the delayed reward, $A$ represents the absolute magnitude of the reward, $D$ is the temporal delay intervening until reward delivery, and $k$ is an empirical parameter indexing the individual’s rate of discounting (impulsivity). Because this function is hyperbolic rather than exponential, the curve displays extreme asymmetry: subjective value plummets precipitously as delay is initially introduced, but levels out as delays become distant. This hyperbolic decay curve—a direct mathematical derivative of matching law formulations—provides the definitive mathematical explanation for dynamic preference reversals, self-control failures, and the neurocomputational mechanics of impulsive pathology.
11. Applied Behavior Analysis and Translational Applications
11.1 Understanding Problem Behavior through the Matching Law
The mathematical principles formulated in Herrnstein’s laboratory provided a revolutionary clinical framework for Applied Behavior Analysis (ABA). Historically, severe aberrant behaviors—such as self-injurious behavior (SIB), physical aggression, property destruction, and explosive tantrums in individuals with developmental disabilities—were frequently conceptualized as isolated clinical pathologies requiring punitive suppression. However, translational behavior analysts, led by researchers such as Wayne Fisher, Brian Iwata, and Timothy Vollmer, recognized that aberrant behavior does not occur in an operant vacuum; it functions as one component of a continuous, concurrent choice matrix.
When an individual with severe autism engages in high-rate self-injurious head-banging, that self-injury ($B_1$) is maintained by environmental reinforcers ($R_1$), which may include immediate caregiver attention, sensory stimulation, or escape from demanding academic tasks. Concurrently, the individual possesses socially appropriate replacement behaviors ($B_2$), such as functional speech, sign language, or picture exchange communication cards, which yield alternative reinforcers ($R_2$). According to the matching law, the proportion of time the individual commits to self-injury is not a fixed attribute of their clinical diagnosis; it is a direct mathematical function of the relative reinforcement rate delivered for self-injury compared to the relative reinforcement rate accessible for adaptive behaviors.
Furthermore, the matching law illuminates why traditional clinical interventions historically suffered from devastating treatment relapses, such as behavioral contrast and resurgence. If a clinician successfully suppresses a target problem behavior in a therapy room by implementing extinction, but fails to enrich reinforcement for appropriate replacement actions, the extraneous reinforcement balance ($R_e$) is fundamentally disrupted. As soon as the client returns to an environment where replacement behaviors go unrecognized and background attention is sparse, the mathematical balance dictated by Herrnstein’s hyperbolic equation forces the problem behavior to resurge with explosive intensity, reclaiming its status as the most efficient conduit to reinforcement.
11.2 Designing Clinical Interventions via Quantitative Choice Architecture
By leveraging the mechanics of the Generalized Matching Law, translational behavior analysts engineered sophisticated, non-aversive intervention architectures, most notably optimized Differential Reinforcement of Alternative Behavior (DRA). Classical DRA approaches often struggled because clinicians simply attempted to reinforce an adaptive behavior while hoping problem behavior would fade. The matching law transformed this practice into a rigorous quantitative science: to eradicate problem behavior, clinicians do not need to deploy intrusive physical restraint or aversive punishments; they merely need to manipulate concurrent schedule parameters so that the relative reinforcement ratio heavily favors the adaptive alternative.
Clinical choice architects achieve this by altering the multi-dimensional facets of reinforcement mapped in Baum’s equations. Clinicians engineer the environment so that the replacement behavior (e.g., handing an exchange card) yields reinforcement that is vastly superior in *immediacy* (zero-second delay versus a long delay for problem behavior), *rate* (continuous reinforcement for communication versus an intermittent schedule for problem behavior), *magnitude* (access to an entire preferred edible versus a small crumb), and *quality* (access to a top-tier sensory toy). Concurrently, by placing the problem behavior on an extinction or thin schedule, its relative reinforcement rate approaches zero.
Even in complex clinical scenarios where dangerous behavior cannot be safely placed on total extinction (e.g., severe eye-gouging that requires immediate physical caregiver intervention, inadvertently maintaining the attention contingency), matching law interventions can successfully eliminate the behavior. By aggressively enriching the background environment with non-contingent reinforcement (NCR)—massively escalating the extraneous reinforcement parameter $R_e$ through free, continuous access to attention, comfort, and leisure items—the relative reinforcing efficacy of the caregiver’s brief, reactive attention following self-injury is mathematically diluted. In accordance with Herrnstein’s hyperbolic equation, inflating $R_e$ drives absolute problem behavior down toward zero, achieving clinical stabilization without aversive contingencies.
11.3 Human Social Interactions and Naturalistic Matching
The explanatory reach of the matching law extends far beyond psychiatric clinics into the rich tapestry of normative human social interaction. In a classic, foundational study conducted in 1974, James Conger and Raymond Killeen evaluated whether conversational dynamics among university students conformed to matching predictions. Unsuspecting human participants engaged in structured group discussions around a table alongside two confederate experimenters. Unknown to the participant, the two confederates delivered verbal praise, head nods, and agreement statements (“Mm-hmm,” “Good point”) according to independent, concurrent variable-interval schedules.
Conger and Killeen tracked the participant’s continuous visual gaze direction and total talking time allocated toward each confederate. The empirical results were astounding: human participants matched their relative gaze duration and verbal conversational time in exact proportion to the relative frequency of social reinforcement delivered by the respective confederates. Without any conscious awareness of the experimental manipulation, the humans behaved identically to Herrnstein’s pigeons, continuously shifting their social investment to match the relative density of social praise, demonstrating that the matching law operates as an unconscious organizing principle of human social discourse.
Subsequent naturalistic studies identified robust matching dynamics across diverse cultural and athletic domains. In professional athletics, researchers demonstrated that the distribution of play-calling in American football (running plays versus passing plays) and shot selection in professional basketball (two-point field goals versus three-point field goals) match the relative points earned per possession across those respective alternatives. In the arena of substance abuse and addiction psychiatry, addiction is increasingly conceptualized through the matching law as a chronic behavioral allocation disorder: an individual’s escalation into drug dependency reflects a concurrent choice matrix where the immediacy, magnitude, and reliability of chemical pharmacological reinforcement overwhelmingly out-competes the delayed, uncertain, and degraded reinforcement available from alternative social, vocational, and familial sources.
12. Neurobiological Foundations, Modern Critiques, and Contemporary Status
12.1 Neural Substrates of Matching and Melioration
The dawn of 21st-century neuroeconomics and cognitive neuroscience initiated an intensive search for the physical neural substrates that compute matching law parameters. Electrophysiological investigations spearheaded by neuroscientists such as Paul Glimcher, Michael Platt, and William Newsome demonstrated that neurons within the mammalian brain actively encode the exact mathematical variables formalized in Herrnstein’s equations. Recording from single neurons within the posterior parietal cortex (specifically the lateral intraparietal area, or LIP) and the frontal eye fields of rhesus macaques engaged in concurrent visual saccade tasks revealed that LIP firing rates systematically scale with the relative reinforcement value—compounding probability, magnitude, and immediacy—of the corresponding movement targets.
At the center of this neural valuation network lies the mesolimbic and nigrostriatal dopaminergic system. Dopamine neurons originating in the ventral tegmental area (VTA) and substantia nigra pars compacta (SNc) project heavily to the nucleus accumbens, ventral striatum, and orbitofrontal cortex (OFC). Groundbreaking computational models developed by Wolfram Schultz, Peter Dayan, and P. Read Montague demonstrated that phasic dopamine bursts do not encode absolute pleasure; instead, they encode Reward Prediction Errors (RPEs)—the mathematical delta between obtained and expected reward. In concurrent schedule tasks, corticostriatal synaptic plasticity continuously updates the subjective value of competing action representations based on these dopaminergic prediction signals.
Remarkably, modern computational neuroscience models, such as temporal difference reinforcement learning (TDRL) and actor-critic networks, demonstrate that when two competing neural populations in the striatum update their synaptic weights via dopamine-mediated prediction errors, the network naturally converges on steady-state firing patterns that mirror melioration. The orbitofrontal cortex evaluates the relative economic goods and updates the subjective valuation matrix, while the basal ganglia circuit arbitrates motor selection. These findings definitively establish that the matching law is not an arbitrary behavioral abstraction, but the macroscopic behavioral manifestation of conserved neurocomputational algorithms etched into the corticostriatal circuitry of the mammalian brain.
12.2 Methodological and Conceptual Critiques of the Matching Law
Despite its monumental empirical success, the matching law has faced sustained methodological and philosophical critiques throughout its history. One persistent conceptual critique centers on the charge of tautology. Skeptics have argued that if reinforcement is defined purely post-hoc as any environmental event that increases the probability of an operant response, and the matching law subsequently states that response rates match reinforcement rates, the equation borders on an unfalsifiable algebraic circularity. If a researcher can arbitrarily designate any unmeasured internal or external stimulus as part of the extraneous reinforcement pool ($R_e$) whenever empirical data deviate from matching, the model risks becoming an untestable tautology rather than a predictive physical law.
A second major methodological challenge concerns the definition and standardization of the operant response unit. In controlled laboratory experiments, a mechanical switch defines a discrete response with unambiguous clarity. However, in complex human and ecological environments, behavior occurs in continuous, overlapping, and fluid topological streams. Determining where one behavioral unit terminates and another begins—and ensuring that competing response topographies carry equivalent biomechanical and metabolic costs—presents immense analytical hurdles. If Manipulandum 1 requires twice the physical caloric energy of Manipulandum 2, treating their raw frequencies as equivalent behavioral units distorts the mathematical integrity of the matching equation.
Furthermore, standard matching law models have faced criticism for their historical difficulty in accommodating profound cognitive anomalies documented by behavioral economists like Daniel Kahneman and Amos Tversky. Phenomena such as loss aversion, the framing effect, asymmetric risk preferences, and the role of complex rule-governed cognitive heuristics often produce marked departures from simple power-function matching. While extended multi-parameter formulations have attempted to incorporate prospect-theoretic value functions into operant equations, critics argue that the matching law’s reliance on stationary steady-state baselines limits its capacity to model rapid, one-shot heuristics executed in volatile, non-equilibrium environments.
12.3 Contemporary Standing and Legacy in 21st Century Behavioral Science
Today, more than six decades after the publication of Herrnstein’s 1961 monograph, the matching law stands as one of the most durable, influential, and thoroughly replicated theoretical achievements in the history of behavioral psychology. By elevating behavior analysis from qualitative descriptions of cumulative curves to the mathematical elegance of quantitative power laws, Richard Herrnstein fundamentally inaugurated the field of quantitative behavior analysis. His work definitively proved that non-human organisms and human agents do not operate through chaotic whim; their choices conform to predictive mathematical equations of extraordinary precision.
In contemporary 21st-century behavioral science, the Generalized Matching Law continues to flourish, integrated into the algorithmic architecture of artificial intelligence, computational psychiatry, and modern digital design. Data scientists and software engineers utilize matching-law equations to model and optimize user engagement within algorithmic recommender systems, designing digital platforms and gamification mechanics that manage user attention budgets across competing concurrent information streams. In clinical neuropsychiatry, computational matching models provide an objective quantitative diagnostic framework for mapping how neural circuit disruptions express themselves as behavioral allocation pathologies across addictive, depressive, and compulsive disorders.
Ultimately, Richard Herrnstein’s enduring legacy is the establishment of a unified, relativistic science of choice. He demonstrated that no action can be understood in isolation; all behavior is fundamentally an allocation problem occurring against a dynamic canvas of competing biological possibilities. By capturing this fundamental truth in an enduring algebraic formulation, Herrnstein permanently bridged the divide separating behavior analysis, microeconomics, and computational neuroscience, cementing the Matching Law as an unshakeable cornerstone in the scientific understanding of biological decision-making.
Conclusion
Richard Herrnstein’s 1961 seminal investigation on concurrent variable-interval schedules forever altered the landscape of psychological science. Prior to his work, the behavioral analysis of operant conditioning remained largely confined to idiographic, single-response baselines that obscured the dynamic, competitive reality of decision-making. By designing an experimental apparatus where pigeons navigated simultaneous, independent reinforcement schedules across dual operant manipulanda, Herrnstein captured the continuous flow of choice in a free-operant laboratory environment. The resulting formulation—the Simple Matching Law—established that an organism’s relative behavioral output directly mirrors the relative reinforcement input yielded by the environment.
Over the ensuing decades, this foundational insight underwent rigorous empirical testing, theoretical expansion, and methodological refinement. William Baum’s Generalized Matching Law expanded the formulation into a robust power function capable of quantifying the distinct parameters of environmental sensitivity and choice bias, providing a psychophysical bridge to classical sensory scaling. Theoretical debates surrounding molecular maximization, molar maximization, and Herrnstein and Prelec’s theory of melioration unmasked the underlying mechanisms of choice, proving that matching is the inexorable equilibrium of local reinforcement rate equalization, operating autonomously from global utility maximization.
Today, the matching law transcends its origins in avian operant chambers. It serves as an indispensable theoretical architecture within behavioral economics, explaining complex intertemporal choice, consumer elasticity, and hyperbolic temporal discounting. In applied clinical arenas, it guides non-aversive interventions for severe behavioral disorders, demonstrating that problem behavior can be treated by mathematically restructuring environmental reinforcement budgets. Simultaneously, contemporary neuroeconomics continues to confirm that dopamine-mediated prediction errors within corticostriatal circuits actively implement the computational algorithms of matching. Herrnstein’s legacy remains an enduring testament to the power of quantitative behavioral science, affirming that within the seeming complexity of biological choice lies a deep, lawful, and mathematically elegant order.
References
- Ainslie, G. (1975). Specious reward: A behavioral theory of impulsiveness and impulse control. Psychological Bulletin, 82(4), 463–496. https://doi.org/10.1037/h0076860
- Baum, W. M. (1974). On two types of deviation from the matching law: Bias and undermatching. Journal of the Experimental Analysis of Behavior, 22(1), 231–242. https://doi.org/10.1901/jeab.1974.22-231
- Baum, W. M. (1979). Matching, undermatching, and overmatching in studies of choice. Journal of the Experimental Analysis of Behavior, 32(2), 269–281. https://doi.org/10.1901/jeab.1979.32-269
- Catania, A. C. (1966). Concurrent operants. In W. K. Honig (Ed.), Operant Behavior: Areas of Research and Application (pp. 213–270). Appleton-Century-Crofts.
- Conger, R., & Killeen, P. (1974). Use of concurrent operants in analyzing talking behavior in humans. Pacific Sociological Review, 17(4), 399–416. https://doi.org/10.2307/1388691
- Davison, M., & McCarthy, D. (1988). The Matching Law: A Research Review. Lawrence Erlbaum Associates.
- Fantino, E. (1969). Choice and rate of reinforcement. Journal of the Experimental Analysis of Behavior, 12(5), 723–730. https://doi.org/10.1901/jeab.1969.12-723
- Ferster, C. B., & Skinner, B. F. (1957). Schedules of Reinforcement. Appleton-Century-Crofts. https://doi.org/10.1037/10627-000
- Fisher, W. W., & Mazur, J. E. (1997). Basic and applied research on choice responding. Journal of Applied Behavior Analysis, 30(3), 387–410. https://doi.org/10.1901/jaba.1997.30-387
- Fleshler, M., & Hoffman, H. S. (1962). A progression for generating variable-interval schedules. Journal of the Experimental Analysis of Behavior, 5(4), 529–530. https://doi.org/10.1901/jeab.1962.5-529
- Glimcher, P. W. (2003). Decisions, Uncertainty, and the Brain: The Science of Neuroeconomics. MIT Press. https://doi.org/10.7551/mitpress/2302.001.0001
- Herrnstein, R. J. (1961). Relative and absolute strength of response as a function of frequency of reinforcement. Journal of the Experimental Analysis of Behavior, 4(3), 267–272. https://doi.org/10.1901/jeab.1961.4-267
- Herrnstein, R. J. (1970). On the law of effect. Journal of the Experimental Analysis of Behavior, 13(2), 243–266. https://doi.org/10.1901/jeab.1970.13-243
- Herrnstein, R. J., & Prelec, D. (1991). Melioration: A theory of distributed choice. Journal of Economic Perspectives, 5(3), 137–156. https://doi.org/10.1257/jep.5.3.137
- Hursh, S. R. (1980). Economic concepts for the analysis of behavior. Journal of the Experimental Analysis of Behavior, 34(2), 219–238. https://doi.org/10.1901/jeab.1980.34-219
- Platt, M. L., & Glimcher, P. W. (1999). Neural correlates of decision variables in parietal cortex. Nature, 400(6741), 233–238. https://doi.org/10.1038/22268
- Rachlin, H., & Green, L. (1972). Commitment, choice and self-control. Journal of the Experimental Analysis of Behavior, 17(1), 15–22. https://doi.org/10.1901/jeab.1972.17-15
- Shimp, C. P. (1966). Probabilistically reinforced choice behavior in pigeons. Journal of the Experimental Analysis of Behavior, 9(4), 443–455. https://doi.org/10.1901/jeab.1966.9-443
- Skinner, B. F. (1938). The Behavior of Organisms: An Experimental Analysis. Appleton-Century-Crofts.
- Vollmer, T. R., & Bourret, J. (2000). Application of the matching law to evaluate the allocation of substance abuse treatment. Journal of Applied Behavior Analysis, 33(4), 433–442. https://doi.org/10.1901/jaba.2000.33-433