Behavioral EconomicsGame Theory

Guessing Game – Rosemarie Nagel The Centipede Game Experiment – Richard

A comprehensive analysis of Rosemarie Nagel’s guessing game and Richard Rosenthal’s centipede game, exploring bounded rationality and backward induction.

memjavad
PUBLISHED
Scientifically Reviewed · Dr. Marwa Abd-Alazim · September 12, 2026
Medically & Scientifically Reviewed Verified: September 12, 2026
Dr. Marwa Abd-Alazim Ph.D.
Professor of Psychology University of Kerbala
Review Criteria & Clinical Standards

This content undergoes rigorous scientific peer-review and medical editorial standards at Arab Psychology Network to ensure clinical accuracy, validity, and compliance with evidence-based guidelines from leading psychological and healthcare authorities (APA / WHO).

The architecture of neoclassical economics has long rested upon the foundational premise of homo economicus: an idealized decision-maker endowed with infinite cognitive capacity, infallible deductive prowess, and an unyielding commitment to personal utility maximization. Within non-cooperative game theory, this axiomatic framework crystallized into the concept of Nash equilibrium and its extensive-form refinements. These models assume not merely that actors act rationally, but that this rationality is common knowledge among all participants. Under common knowledge of rationality, each agent knows that all other agents are rational, knows that everyone knows it, and can execute an infinite chain of nested logical deductions without friction, error, or psychological hesitation. For decades, this mathematical paradigm offered elegant, analytically tractable solutions to strategic interactions ranging from market competition to geopolitical conflicts.

However, the pristine abstractions of pure equilibrium analysis inevitably encountered friction when confronted with empirical human behavior. When real human beings enter laboratory environments or navigate real-world markets, the clean predictions of classical game theory often fail to materialize. Instead of frictionless convergence to theoretical equilibria, researchers consistently observe pervasive, systematic, and structurally reproducible departures from normative predictions. These behavioral regularities do not represent mere chaotic noise or cognitive defects; rather, they reflect the intrinsic cognitive architecture of human decision-makers operating under finite processing capabilities, endogenous computational bounds, and nuanced social motivations. The divergence between mathematical elegance and empirical reality necessitated a fundamental paradigm shift: the emergence of behavioral game theory.

At the vanguard of this empirical revolution stand two foundational experimental paradigms: Rosemarie Nagel’s simultaneous-move Guessing Game (widely known as the p-beauty contest) and Richard Rosenthal’s sequential-move Centipede Game. Though fundamentally distinct in their strategic topology—one operating in normal form, the other in extensive form—both experimental frameworks systematically exposed the limits of classical deduction. Nagel’s work illuminated the finite depths of iterative strategic reasoning through which human minds forecast the choices of others, while Rosenthal’s formulation brought to light the dramatic behavioral and epistemic paradoxes inherent in backward induction. Together, these two paradigms altered modern economics, transforming game theory from a purely normative mathematical exercise into an empirical social science grounded in the reality of human cognition.

1. Introduction to Experimental Game Theory and Bounded Rationality

1.1 Historical Emergence of Behavioral Game Theory

The development of behavioral game theory emerged as an intellectual reaction against the axiomatic rigidities that characterized mid-twentieth-century mathematical economics. Following the seminal contributions of John von Neumann, Oskar Morgenstern, and John Nash, game theory established itself as the premier mathematical language for interactive decision-making. However, early theorists operated almost exclusively within a normative tradition. They sought to identify what hyper-rational agents ought to do when facing equally omniscient counterparts, largely dismissing empirical testing as irrelevant to pure logic. Models assumed that players possessed unlimited computational bandwidth, free from psychological fatigue, subjective biases, or cognitive heuristics. This detached mathematical architecture insulated the discipline from empirical validation for decades.

The transformation began as experimental economics gained institutional and methodological legitimacy through the pioneering work of figures such as Vernon Smith, Reinhard Selten, and Amos Tversky and Daniel Kahneman. Researchers began to recognize that economic theories, like any empirical scientific claims, required rigorous laboratory testing under controlled, reproducible conditions. The early pioneers established strict methodological canons: the deployment of salient financial incentives to motivate authentic decision-making, the elimination of deceptive practices to maintain experimental control, and the creation of neutral, context-free protocols to isolate strategic calculations from idiosyncratic real-world associations. Through these methods, economists could systematically observe how flesh-and-blood individuals actually behave within rigorously defined strategic structures.

Within this emerging landscape, Rosemarie Nagel and Richard Rosenthal played indispensable roles in exposing the phenomenon of bounded rationality, a concept first articulated by Herbert Simon. In the mid-1990s, Nagel introduced the experimental p-beauty contest game, providing the discipline with an exceptionally precise instrument to calibrate and measure the exact depth of iterative human reasoning in simultaneous-move environments. Concurrently, Rosenthal’s 1981 theoretical formulation of the Centipede Game—subsequently tested experimentally by Richard McKelvey and Thomas Palfrey in 1992—presented an empirical paradox: a game where the unique, mathematically ironclad backward induction equilibrium was routinely and spectacularly violated by human players. These foundational contributions firmly demonstrated that human strategic choice is not arbitrary, but rather governed by systematic, mathematically modelable cognitive constraints.

1.2 Dichotomy Between Equilibrium Predictions and Human Decision-Making

The core tension between traditional game theory and behavioral observation lies in the distinction between mathematical tractability and human psychological architecture. Neoclassical models demand that players form mutually consistent expectations. In a Nash equilibrium, every player’s strategy must be an optimal response to the actual strategies chosen by everyone else. To achieve this state purely through deduction, players must be capable of processing infinite regresses of belief: “I believe that you believe that I believe…” ad infinitum. While this condition provides mathematical closure, it assumes away the formidable cognitive burdens that real minds experience when attempting to track multi-layered recursive mental states.

In actual strategic environments, human agents encounter severe cognitive limitations, including constraints on working memory capacity, computational processing speed, and the executive control required to project recursive mental models. Consequently, the assumption of common knowledge of rationality collapses. A decision-maker may be entirely rational in isolation, yet if they harbor doubts about whether their counterpart possesses identical rationality—or if they suspect their counterpart doubts their rationality—the normative justification for equilibrium play completely breaks down. Rational belief formation under real-world conditions requires players to formulate subjective probability distributions over how others will act, incorporating the likelihood of cognitive errors, idiosyncratic deviations, and bounded depths of calculation.

This dynamic introduces the critical problem of out-of-equilibrium behavior. In classical game theory, events that occur with zero probability in equilibrium are essentially undefined; the theory offers limited guidance on how a rational player should interpret an empirical move that violates the equilibrium path. Yet, in laboratory and practical environments, out-of-equilibrium play is the norm rather than the exception. When an opponent deviates from the prescribed equilibrium trajectory, a human player must engage in immediate cognitive attribution: Was the deviation an accidental trembling-hand error, an act of calculated altruism, an attempt to signal future cooperation, or a manifestation of profound strategic ignorance? The divergence between the mathematical fixed points of equilibrium analysis and the dynamic, belief-updating processes of real human beings defines the foundational inquiry of modern behavioral economics.

1.3 Methodological Scope: Normal Form Versus Extensive Form Experimental Designs

To systematically map the boundaries of human strategic reasoning, experimental game theorists partition interactive environments into two primary structural archetypes: normal (strategic) form games and extensive form games. These two structures impose fundamentally different cognitive demands on participants, engaging distinct psychological mechanisms and computational pathways. Investigating both formats allows researchers to develop a robust, multi-dimensional taxonomy of bounded rationality across varying institutional settings.

Normal form games, exemplified by Rosemarie Nagel’s Guessing Game, represent strategic scenarios as simultaneous interactions. Players must make their decisions in isolation, without observing the contemporaneous choices of their peers. From a cognitive perspective, normal form games compel agents to engage in internal, prospective simulation. A player must construct an internal model of the entire population, infer the probabilistic distribution of strategies that others will execute, and calculate an optimal response to those aggregated expectations. There is no historical trace, no dynamic feedback during the decision moment, and no opportunity to adjust one’s strategy mid-game. The cognitive load centers on the simultaneous synthesis of mutual expectations and parallel depth of reasoning.

Conversely, extensive form games, epitomized by Richard Rosenthal’s Centipede Game, introduce dynamic, sequential mechanics unfurling over time across discrete decision nodes. In these environments, players observe the explicit, historical choices made by their counterparts at prior nodes before selecting their own actions. The cognitive operation required is not parallel expectation aggregation, but rather sequential temporal forecasting via backward induction. Players are theoretically required to mentally leap to the absolute end of the decision tree, deduce the terminal payoffs, identify the optimal action at the final node, and then fold the game tree backward step by step to determine the rational move at the present moment. By comparing how human subjects navigate the simultaneous demands of the Guessing Game alongside the sequential unraveling of the Centipede Game, experimental economists can rigorously delineate the universal cognitive constraints that govern human interaction across diverse strategic topologies.

2. Theoretical Framework of Rosemarie Nagel’s Guessing Game (p-Beauty Contest)

2.1 John Maynard Keynes’ Newspaper Beauty Contest Metaphor

The intellectual lineage of the modern Guessing Game traces directly back to chapter 12 of John Maynard Keynes’s masterwork, The General Theory of Employment, Interest and Money (1936). In analyzing the speculative dynamics of financial markets, Keynes observed that professional investment closely resembles a popular newspaper competition of his era. In these contests, readers were presented with a grid of one hundred photographs of attractive faces and tasked with selecting the six most beautiful. The prize was awarded not to the person who selected the faces they personally found most aesthetically pleasing, nor to the person who chose the faces deemed objectively most beautiful by some absolute standard, but to the contestant whose selections most closely mirrored the average preferences of all competitors combined.

Keynes recognized that this structure radically alters the nature of the cognitive task. An introspective, non-strategic participant—operating at what modern theorists classify as a naive level—merely chooses the faces they personally find most attractive. A more sophisticated contestant, however, immediately realizes that personal taste is irrelevant; the goal is to forecast aggregate taste. Therefore, this second-order player selects the faces that they believe the average reader will find attractive. Yet this line of reasoning cannot stop there. If other contestants are equally sophisticated, they too will anticipate this logic. Consequently, third-order reasoning emerges: contestants attempt to anticipate what average opinion expects average opinion to be. Keynes famously noted that some individuals advance to the fourth, fifth, and higher degrees of iterative speculation, though he presciently recognized that human cognitive stamina rapidly exhausts itself within this recursive hall of mirrors.

This newspaper contest provides a profound metaphor for the pricing of liquid assets in equity and capital markets. Keynes argued that stock valuations rarely reflect fundamental, intrinsic cash-flow projections calculated by rational analysts. Instead, market participants spend their energies attempting to anticipate what the market consensus will believe the consensus will be three months or a year into the future. The game is one of pure strategic complementarity: the payoff to an action depends directly on how many other actors coordinate on that exact same action. The higher-order nature of mutual belief estimation decouples asset prices from underlying fundamentals, creating structural vulnerabilities that manifest as speculative manias, valuation bubbles, and sudden liquidity panics.

2.2 Rules, Payoff Structures, and Mechanics of the Nagel Protocol

In her groundbreaking 1995 paper published in The American Economic Review, Rosemarie Nagel translated Keynes’s literary metaphor into an exact, mathematically rigorous, and empirically testable laboratory protocol: the p-beauty contest game. In the baseline version of Nagel’s design, a group of N players (typically ranging from 15 to 20 participants in standard laboratory environments) are gathered and instructed to simultaneously and independently submit a single real number from a closed interval, traditionally defined as:

Si ∈ [0, 100]

The rules of the game are made common knowledge: all instructions are read aloud, and comprehension checks are administered to guarantee that every subject understands the rules and payoff structure. The winning condition is mathematically defined: once all numbers are collected, the experimenter calculates the arithmetic mean of all submitted entries, denoted as M. This mean is then multiplied by a publicly known scaling parameter, p, where 0 < p < 1 (most commonly selected as p = 2/3 or p = 1/2). The target value, T, is thus explicitly computed as:

T = p × M = p × (1 / N) ∑j=1…N xj

The player whose submitted number is closest in absolute distance to the target value T wins a fixed monetary prize. In the event of a tie where multiple players submit numbers equidistant from the target, the prize is divided equally among the winners. All other players receive zero payoff for that round. This winner-take-all incentive structure imposes extreme strategic sharpness: there are no points awarded for simply being close in an absolute sense; one’s utility is strictly contingent on predicting the aggregate distribution of choices more accurately than any competitor in the room. The mechanism elegantly captures the tension between unconstrained individual choice and the collective gravity of the aggregate mean scaled by the parameter p.

2.3 Analytical Determination of the Unique Nash Equilibrium

From the vantage point of normative game theory, the p-beauty contest with p = 2/3 possesses a unique, mathematically inescapable Nash equilibrium. This equilibrium can be derived through the systematic, step-by-step application of the Iterated Elimination of Strictly (and Weakly) Dominated Strategies (IEDS). Because this analytical process requires zero empirical input beyond the formal rules and the assumption of common knowledge of rationality, it serves as the classical benchmark against which human behavior is measured.

The deductive proof proceeds through an infinite sequence of iterative bounds:

  • Step 1: Consider the maximum possible average that the group could theoretically produce. Even if every single participant in the room were to coordinate on the theoretical upper boundary of the strategy space and submit 100, the resulting mean M would be precisely 100. Consequently, the maximum possible target value would be:

    Tmax = 2/3 × 100 = 66.67

    Any rational player who understands this arithmetic constraint realizes immediately that choosing any number in the open interval (66.67, 100] is strictly dominated. Under no conceivable configuration of choices could a number above 66.67 win the game, because the target can never mathematically exceed that value. Therefore, a rational player eliminates all strategies greater than 66.67 from their consideration set.

  • Step 2: If it is common knowledge that all players are rational, every participant knows that no one will choose a number greater than 66.67. Thus, the effective strategy space for all players contracts to the interval [0, 66.67]. Now, the iterative deduction repeats. If the maximum number anyone will choose is 66.67, then the highest conceivable mean the group can generate is 66.67. It follows that the new upper bound for the target value must be:

    T ≤ 2/3 × 66.67 = 44.44

    Consequently, all strategies lying in the interval (44.44, 66.67] are now recognized as dominated and must be discarded.

  • Step 3 through k: This logical elimination cascades recursively through an infinite regress:

    Sk = 100 × (2/3)k

    As the depth of iterative deduction k approaches infinity (k → ∞), the term (2/3)k converges asymptotically to 0.

The only strategy that survives the infinite iteration of dominated strategies is the number 0. Furthermore, the profile where every player chooses 0 constitutes the unique Nash equilibrium of the game: if everyone chooses 0, the mean is 0, the target is 2/3 × 0 = 0, and no individual player has any unilateral incentive to deviate to a higher number. Yet, this proof relies upon an extraordinary epistemic demand: that every single player is rational, knows that all others are rational, knows that everyone knows it, and carries out this transfinite deduction instantaneously.

3. Empirical Findings and Behavioral Distributions in Nagel’s Experiments

3.1 Distribution of Initial Round Strategic Choices

When Rosemarie Nagel administered this game to actual human subjects in laboratory conditions, the empirical results directly shattered the normative predictions of classical equilibrium analysis. In the initial, first-period round of play—where subjects made their choices based purely on introspection without having seen any historical data or feedback—the number of participants who chose the theoretical Nash equilibrium of 0 was virtually non-existent, consistently accounting for less than one percent of the subject pool.

Rather than observing a smooth uniform spread or a sharp spike at zero, Nagel observed a striking empirical regularity that has since been replicated across dozens of studies worldwide: the choices clustered systematically around distinct, highly defined modal peaks. Specifically, when p = 2/3, large concentrations of choices emerged precisely around:

  • 50: Representing a completely non-strategic or central anchor.
  • 33.33: Mathematically corresponding to 2/3 of 50.
  • 22.22: Mathematically corresponding to 2/3 of 33.33 (or (2/3)2 of 50).
  • 15: Corresponding approximately to 2/3 of 22.22 (or (2/3)3 of 50).

The existence of these discrete behavioral spikes provided decisive evidence against two prevailing competing theories: it proved that human players were neither behaving randomly (which would have yielded an even, flat distribution across the [0, 100] interval) nor behaving as hyper-rational equilibrium maximizers (which would have collapsed the entire distribution to 0). Instead, the data revealed an underlying psychological order—a structured, step-by-step heuristic process that was universally shared across human populations, yet fundamentally bounded in its depth.

3.2 Formalization of the Level-k Reasoning Architecture

To provide a rigorous mathematical framework capable of explaining the observed empirical clusters, Rosemarie Nagel developed what is now celebrated as the Level-k model of cognitive hierarchy. This model formalizes the insight that individuals differ systematically in the depth of strategic recursion they deploy when forecasting the behavior of others. Rather than assuming common knowledge of rationality, the model organizes the population into a structured hierarchy of cognitive types, indexed by the integer k, where k represents the exact number of iterative deductive steps an individual executes.

The hierarchy is formally anchored at the base level, known as Level-0 (L0). A Level-0 player is defined as a non-strategic actor who does not engage in any interactive thinking. In the standard specification, a Level-0 agent is assumed to randomize uniformly across the entire available strategy space, selecting any number in the interval [0, 100] with equal probability. Consequently, the mathematical expectation of a Level-0 player’s choice is simply the midpoint of the domain:

E[SL0] = (0 + 100) / 2 = 50

Ascending the cognitive hierarchy, higher-level players are modeled as best-responding to the perceived behavior of players below them:

  • Level-1 (L1): An L1 player assumes that all other competitors in the game are non-strategic Level-0 actors. Therefore, the L1 player expects the aggregate mean to be 50. In order to win, they best-respond by multiplying this expected mean by the parameter p:

    SL1 = p × 50 = (2/3) × 50 = 33.33

    This precisely explained the massive empirical spike observed at 33 in Nagel’s experimental data.

  • Level-2 (L2): An L2 player operates under the model that other participants are Level-1 thinkers. They anticipate that the population will submit choices clustering around 33.33. Consequently, the L2 player calculates their optimal choice as:

    SL2 = p × SL1 = p2 × 50 = (2/3)2 × 50 = 22.22

    This accounted for the second massive empirical peak in the laboratory data.

  • Level-3 (L3): Taking this recursive logic one step further, a Level-3 player assumes the market consists of Level-2 agents and computes:

    SL3 = p × SL2 = p3 × 50 = (2/3)3 × 50 ≈ 14.81

Through structural econometric estimation across extensive laboratory trials, Nagel and subsequent researchers demonstrated that human populations are overwhelmingly dominated by Level-1 and Level-2 thinkers. The mean depth of strategic reasoning in unstudied human cohorts hovers reliably between 1.4 and 1.8 steps. Extremely few individuals naturally execute three or more steps of reasoning in initial interactions, and those who possess the theoretical sophistication to calculate the Nash equilibrium (0) often score poorly in single-shot games because choosing 0 constitutes a catastrophic miscalculation when playing against an empirical population of Level-1 and Level-2 actors. Success in strategic environments requires not pure mathematical deduction, but an accurate psychological appraisal of the bounded reasoning of others.

3.3 Dynamic Learning and Multi-Period Convergence to Equilibrium

While classical equilibrium analysis utterly fails to describe first-period human behavior, Nagel’s experimental protocol yielded an equally profound discovery when the game was repeated over multiple consecutive rounds. In these multi-period treatments, subjects played the guessing game across four to ten consecutive periods. At the conclusion of each round, the experimenter announced the actual mean, the calculated target value, and the winning number, allowing all participants to update their beliefs before submitting their next choice.

Under the influence of this dynamic feedback, the behavioral distribution exhibited a rapid, systematic unraveling toward the theoretical Nash equilibrium. In period two, the choices that had previously clustered around 33 and 22 collapsed downward toward 15 and 10. By period four or five, the aggregate distribution was heavily compressed within the single digits, with significant proportions of the subject pool actively submitting 0 or numbers infinitesimally close to it. The system demonstrated clear empirical convergence toward the fixed point predicted by normative theory.

Crucially, however, Nagel’s analysis revealed that this convergence was driven by directional learning theory rather than an instantaneous intellectual realization of common knowledge of rationality. Players adjusted their choices through backward-looking adaptive heuristics: they observed the target from the previous round and adjusted their next bid downward by applying their characteristic depth of reasoning to the newly established reference point. The cognitive hierarchy remained remarkably stable; what changed was the baseline upon which that hierarchy operated. Rather than transforming into hyper-rational agents, players simply anchored their finite 1- or 2-step iterative calculations onto the observed outcome of the preceding period, dragging the aggregate mean down to zero through successive rounds of experiential adaptation.

4. Cognitive Hierarchy Models and Extensions of the Guessing Game

4.1 The Camerer, Ho, and Chong Cognitive Hierarchy Formulation

While Rosemarie Nagel’s foundational Level-k model revolutionized strategic analysis, it maintained a somewhat restrictive assumption: each Level-k player believed with absolute certainty that all other players in the game belonged exclusively to the single level directly beneath them (Level k − 1). To address this psychological implausibility and provide greater econometric flexibility, Colin Camerer, Teck-Hua Ho, and Kuan-Chong Chong (2004) developed the generalized Cognitive Hierarchy (CH) model.

In the Camerer-Ho-Chong formulation, players recognize that they inhabit a heterogeneous world populated by actors possessing diverse depths of strategic capability. A player of thinking level k (where k ∈ {0, 1, 2, …}) understands that other participants are distributed across all thinking levels from 0 up to k − 1. However, due to an intrinsic cognitive blind spot known as self-referential boundedness, a Level-k player cannot conceive of the existence of players possessing reasoning levels equal to or greater than their own; they view themselves as occupying the apex of strategic sophistication.

Camerer, Ho, and Chong discovered that the empirical distribution of cognitive types across an astonishingly broad array of human societies can be parsimoniously characterized by a single-parameter Poisson distribution:

f(k) = (eτ τk) / k!

In this elegant formulation, the single free parameter τ (tau) represents both the mean and the variance of the depth of reasoning within the population. A Level-k player computes a normalized, truncated conditional probability distribution over the lower-level types they believe exist:

pk(h) = f(h) / ∑l=0…k−1 f(l),    for all h < k

A Level-k player then calculates the expected strategy chosen by each lower type h, aggregates these strategies according to their perceived frequencies, and generates a mathematically precise best-response to that composite mixture. Across dozens of empirical datasets—encompassing college students, portfolio managers, corporate CEOs, and high schoolers—the estimated value of τ reliably hovers around 1.5. This remarkably stable parameter implies that the human species, on average, operates between one and two steps of strategic anticipation, providing economists with a unified, predictive structural model applicable far beyond the laboratory.

4.2 Experimental Variations: Asymmetric Targets, Group Sizes, and Information Sets

Following the widespread recognition of the baseline guessing game, experimental economists devised sophisticated variations to test the resilience, boundary conditions, and behavioral sensitivities of cognitive hierarchies. One of the most theoretically revealing variations involved manipulating the scaling parameter p such that it exceeded unity, setting for example p = 1.3 or p = 1.5 within the choice domain [0, 100].

When p > 1, the directional logic of the iterated elimination of dominated strategies inverts entirely:

  • If p = 1.5, any average choice greater than 0 is magnified.
  • If everyone chooses 100, the target is 150 (bounded at 100).
  • Iterated elimination of dominated strategies cascades upward rather than downward, driving the unique Nash equilibrium to the maximum boundary: 100.

Empirically, when p > 1, human subjects exhibit far faster convergence toward the equilibrium of 100 than they do toward 0 when p < 1. Researchers attribute this asymmetry to psychological framing: human minds process multiplication and upward expansion far more intuitively than fractional decay and recursive division. The cognitive friction of dividing repeatedly by 1.5 or multiplying by 2/3 presents an arithmetic drag that amplifies bounded rationality in fractional games.

Other variations investigated the structural impacts of group size, choice domains, and information structures. Experiments conducted with small groups (e.g., N = 3) revealed profound strategic distortions: in small groups, an individual’s own submitted number exerts a substantial algebraic influence on the resulting mean, transforming the game into a complex exercise in self-referential algebra. Conversely, large-scale macro-experiments executed via national newspapers (such as trials run in the Financial Times, Spektrum der Wissenschaft, and Expansion) involving thousands of concurrent participants demonstrated that even across vast, heterogeneous populations, the modal clusters at 33, 22, and 0 endure with uncanny consistency, confirming that bounded depth of reasoning is an intrinsic feature of human cognition rather than an artifact of small laboratory samples.

4.3 Neuroeconomic and Eye-Tracking Corroboration of Reasoning Depths

In recent years, the behavioral patterns identified by Nagel, Camerer, and their colleagues have received striking physical corroboration through modern neuroeconomics and physiological monitoring technologies. Rather than relying solely on the final numbers written on paper or entered into a computer terminal, researchers can now directly observe the cognitive machinery as it processes strategic information in real time.

Functional Magnetic Resonance Imaging (fMRI) studies conducted during beauty contest games have mapped distinct neural signatures that correlate precisely with estimated Level-k reasoning depths:

  • Participants engaging in basic Level-1 reasoning exhibit elevated metabolic activation primarily in low-level computational and visual processing areas, such as the parietal cortex.
  • Participants executing sophisticated Level-2 and Level-3 reasoning demonstrate intense, localized blood-oxygen-level-dependent (BOLD) signals within the medial prefrontal cortex (mPFC) and the dorsolateral prefrontal cortex (dlPFC). These specific structures are universally recognized as the neurological seats of higher-order executive functioning, mentalizing, and Theory of Mind—the explicit neural capacity to simulate the psychological perspectives and cognitive operations of other human beings.

Simultaneously, high-resolution eye-tracking experiments utilizing computerized matrix interfaces have exposed the visual information-acquisition paths that precede strategic choice. By tracking the exact sequence, duration, and saccadic jumps of a player’s gaze across payoff tables and instruction parameters, researchers can reconstruct the algorithm executed by the subject’s brain. Level-0 and Level-1 subjects visually fixate almost entirely on their own payoff columns and central numbers. In stark contrast, individuals classified as Level-2 demonstrate systematic, recursive eye movements: their gaze darts back and forth between their own potential choices, the target parameters, and the simulated response cells of their opponents. These physiological methodologies provide undeniable biological proof that Level-k models describe authentic mental operations occurring inside the human cranium.

5. Theoretical Foundations of Richard Rosenthal’s Centipede Game

5.1 Formulation and Strategic Topology of the Centipede Game

While Rosemarie Nagel’s guessing game exposed the limits of simultaneous recursive reasoning, an equally profound theoretical crisis was brewing within extensive-form, dynamic game theory. In 1981, Richard Rosenthal published an unassuming two-page paper in the Journal of Mathematical Sociology titled “An Paradox of Backward Induction,” introducing a game that would forever alter the study of dynamic strategic interactions: the Centipede Game.

The Centipede Game is a finite extensive-form game of perfect information played between two actors, conventionally designated as Player 1 and Player 2. The game features an alternating sequential structure that visually resembles a centipede when rendered as an extensive-form game tree. The game begins at Node 1 with Player 1. In front of the players are two piles of money. At any decision node t, the active player faces a binary choice between two discrete actions:

  • Take (T) / Defect: The player terminates the game immediately, seizing the larger share of the accumulated financial piles and allocating the smaller share to their opponent.
  • Pass (P) / Cooperate: The player passes the decision to their counterpart at the subsequent node. Crucially, passing triggers an exogenous multiplication of the total financial pot—every time a player passes, the collective wealth increases.

Mathematically, let the nodes be indexed n = 1, 2, …, M. At any node n, if the active player chooses to Take (terminate), the payoffs are formally structured such that:

  • The terminating player receives: Uactive(n) = f(n) + c
  • The non-terminating player receives: Upassive(n) = f(n) − d

Where f(n) grows monotonically with the node index n. However, if the active player chooses to Pass, the total pot expands such that if the subsequent player were to pass, both players would achieve a payoff vastly superior to the termination values at node n. The game represents a profound dynamic tension: cooperation continuously inflates the collective surplus, yet at every single discrete juncture, the individual possessing the move faces an immediate, irresistible financial incentive to defect and exploit their counterpart’s prior benevolence.

5.2 Backward Induction and Subgame Perfect Nash Equilibrium (SPNE)

Under standard neoclassical game theory, any finite extensive-form game of perfect information can be solved definitively through the application of backward induction, a principle mathematically grounded in the classical Zermelo-von Neumann-Kuhn theorem. To determine the uniquely rational course of action, one does not begin at the beginning; rather, one must mentally transport themselves to the absolute terminal node of the game tree and solve the system in reverse chronological order.

Consider a standard finite Centipede Game designed to terminate conclusively at Node M (assume M is even, making Player 2 the final decision-maker). At Node M, Player 2 faces a completely deterministic choice:

  • If Player 2 chooses Take, they receive a large payoff of, say, $10.00, while Player 1 receives$8.00.
  • If Player 2 chooses Pass, the game concludes exogenously, yielding a payoff where both players receive an equal cooperative sum, say, $9.00 each.

Because Player 2 is an axiomatically rational, selfish utility maximizer, they compare $10.00 to$9.00. Since $10.00 >$9.00, Player 2 will unambiguously choose Take at Node M. There is no uncertainty, no future to protect, and no fear of retaliation.

Now, backward induction steps back to Node M − 1, where Player 1 must select their action. Player 1 possesses complete information and knows with absolute mathematical certainty that if they pass the move to Player 2, Player 2 will strictly choose Take at Node M, leaving Player 1 with $8.00. Alternatively, if Player 1 chooses Take right now at Node M − 1, Player 1 secures an immediate payoff of $8.50, leaving Player 2 with$6.50. Player 1 evaluates their binary alternatives: take now for $8.50, or pass and receive$8.00 at the next node. Because $8.50 >$8.00, Player 1 unambiguously chooses Take at Node M − 1.

This deductive contagion cascades backward through the entire game tree. At Node M − 2, Player 2 anticipates Player 1’s defection at Node M − 1 and chooses to defect at M − 2. The logic unravels inexorably, link by link, all the way back to the very first decision point:

The Subgame Perfect Nash Equilibrium (SPNE) dictates that Player 1 must choose Take at Node 1.

The game terminates instantaneously at the very first step. Player 1 walks away with pennies, Player 2 walks away with nothing, and the vast mountain of potential collective wealth generated across the subsequent nodes is entirely forfeited. Normative game theory delivers an outcome that is uniquely rational, mathematically pristine, and monumentally, breathtakingly inefficient.

5.3 The Epistemic Paradox of Backward Induction

The profound divergence between the theoretical prediction and basic human intuition exposed what philosophers and epistemic game theorists term the paradox of backward induction. The problem lies not merely in the tragic inefficiency of the outcome, but in a profound logical contradiction buried within the epistemic foundations of extensive-form rationality itself, a critique famously articulated by theorists such as Ken Binmore and Philip Reny.

The derivation of Subgame Perfect Nash Equilibrium requires the assumption of Common Knowledge of Rationality (CKR). Every deduction made by a player relies on the absolute premise that all players will make rational choices at all future nodes under all circumstances. Now, consider the philosophical dilemma that confronts Player 2 if they suddenly find themselves called upon to make a move at Node 2. According to the rock-solid mathematics of backward induction, Node 2 is an event that occurs with a probability of exactly zero. It should never be reached in this universe, because Player 1 should have terminated the game instantly at Node 1.

The very fact that Player 2 is holding the move proves empirically that Player 1 has committed an act that directly contradicts the axiomatic model of hyper-rationality. How, then, is Player 2 supposed to interpret this empirical reality? The classical epistemic framework collapses:

  • Should Player 2 assume that Player 1 is an irrational fool who does not understand game theory? If so, Player 1 cannot be relied upon to play the SPNE strategy at Node 3, which completely invalidates the backward induction calculation that justified defecting in the first place!
  • Should Player 2 assume that Player 1 is a strategic mastermind who passed at Node 1 precisely to signal an invitation to mutually lucrative cooperative play?
  • Or was Player 1’s move a momentary lapse of concentration—a random “tremble” in the sense of Reinhard Selten?

Standard equilibrium theory provides no framework for how an agent should revise their conditional beliefs once an axiom of the theory has been explicitly falsified by observed history. The moment a player deviates from the equilibrium path, the assumption of common knowledge of rationality is shattered, and the deductive chain required to justify backward induction evaporates into an epistemic void.

6. Empirical Execution: McKelvey and Palfrey’s Centipede Experiments

6.1 Experimental Design and Laboratory Protocols of McKelvey and Palfrey (1992)

For more than a decade following Rosenthal’s 1981 paper, the Centipede Game remained a theoretical curiosity—a provocative mathematical thought experiment debated in seminars but untested in the physical world. That status changed definitively in 1992 with the publication of Richard McKelvey and Thomas Palfrey’s landmark paper, “An Experimental Study of Several Centipede Games,” published in Econometrica.

McKelvey and Palfrey brought the Centipede Game into the controlled environment of the experimental laboratory, designing an empirical architecture specifically constructed to eliminate confounding real-world artifacts:

  • Structural Variants: They implemented both a four-node and a six-node extensive-form centipede game, allowing them to examine whether tree length influenced the propensity to cooperate.
  • Random Matching Protocol: To systematically prevent the emergence of repeated-game cooperation, reputation effects, or collusive supergame strategies, they utilized a strict random stranger matching protocol. In each trial, subjects were randomly repaired with an anonymous counterpart, ensuring that no pair of individuals ever played each other repeatedly.
  • Real Monetary Stakes: Every choice involved salient, tangible financial rewards paid in cash at the end of the session. Furthermore, to test the durability of the behavior against wealth effects, they executed high-stakes treatments where all payoffs were multiplied by a factor of four.
  • Absolute Anonymity: Stringent double-blind procedures were maintained to eliminate experimenter demand effects and prevent post-experiment social sanctions between participants.

6.2 Empirical Findings: Node-by-Node Departure from Equilibrium

The experimental results obtained by McKelvey and Palfrey dealt a catastrophic blow to the predictive validity of Subgame Perfect Nash Equilibrium. If the classical theory of backward induction were an accurate descriptor of human decision-making, 100% of all games would have terminated decisively at Node 1. Instead, McKelvey and Palfrey observed that:

Immediate first-node defection occurred in less than 7% of all games played.

Rather than collapsing at the starting line, human participants systematically, reliably, and cheerfully passed the move forward, plunging deep into the extensive-form tree. In the four-node games, players routinely cooperated through Nodes 1, 2, and 3, with termination clustering predominantly at Node 3 (Player 1 defecting) and Node 4 (Player 2 defecting). Even more astonishingly, in a non-trivial fraction of trials—approximately 37% of games in certain sessions—the game reached the absolute terminal node, and in some instances, Player 2 actually passed at the final node, granting both players the maximum possible cooperative surplus!

In the six-node games, the empirical distribution shifted outward in tandem with the expanded horizon. Termination choices formed a bell-shaped distribution peaked around Nodes 4 and 5. Rather than exhibiting immediate selfish liquidation, participants generated vast amounts of collective wealth that classical economics declared impossible. The high-stakes treatments, far from forcing subjects to conform to the Subgame Perfect equilibrium, yielded behavioral distributions that were virtually indistinguishable from the low-stakes baseline. The departure from backward induction was not a cheap consequence of trivial incentives, but a deeply ingrained, robust feature of human strategic interaction.

6.3 Learning Dynamics and Repetition Over Time

A central defense often mounted by traditional game theorists is that classical rationality represents an asymptotic state achieved through experiential learning: while naive subjects might err initially, exposure to market forces and repeated exploitation will inevitably discipline them into equilibrium play. To rigorously test this hypothesis, McKelvey and Palfrey had their subjects play the game across multiple successive rounds (typically 10 to 15 iterations) against constantly reshuffled, anonymous counterparts.

The empirical data revealed that experiential learning did indeed take place, but not in the manner predicted by classical theory:

  • Over successive rounds, an undeniable process of backward unraveling occurred. As players experienced the pain of being terminated upon by their counterparts at later nodes, they adapted by preemptively defecting one step earlier in subsequent rounds. The empirical distribution of terminal nodes drifted systematically to the left over time.
  • However, this unraveling did not collapse to Node 1! Even after ten rounds of experience, when players thoroughly understood the strategic hazards of the game, first-node defection remained exceptionally rare, almost never exceeding 15% of total play.
  • Instead of collapsing to the unique Nash equilibrium, the system stabilized into a robust, dynamic behavioral steady-state. Players learned to cooperate across the initial nodes where the mutual gains from pot expansion were massive, while concentrating their preemptive defections in the penultimate stages of the tree. The human capacity to sustain conditional cooperation proved remarkably resilient against complete game-theoretic decay.

7. Explaining the Centipede Paradox: Altruism, Reciprocity, and Social Preferences

7.1 Inequity Aversion and Altruistic Payoff Transformations

Faced with the dramatic empirical failure of standard game theory to predict Centipede play, behavioral economists recognized that the flaw lay not in human intelligence, but in the narrow specification of human objective functions. Neoclassical theory assumed that an agent’s utility is strictly identical to their personal monetary payoff: Ui(x) = xi. In reality, human beings possess rich, complex other-regarding preferences that actively incorporate the payoffs and well-being of their counterparts.

One of the most powerful theoretical frameworks developed to operationalize this insight is the Inequity Aversion model formulated by Ernst Fehr and Klaus Schmidt (1999). Under the Fehr-Schmidt specification, an individual experiences psychological utility from their own financial gain, but suffers an explicit psychological utility penalty whenever an interaction generates an unequal distribution of outcomes. The utility function for Player i facing an outcome vector x is defined as:

Ui(x) = xiαi max{xjxi, 0} − βi max{xixj, 0}

In this equation:

  • The parameter αi measures the strength of Player i‘s disadvantageous inequity aversion (envy): the psychological pain felt when someone else receives more than them.
  • The parameter βi measures the strength of their advantageous inequity aversion (guilt): the moral discomfort felt when they receive more than their counterpart (with 0 ≤ βi < 1 and αiβi).

When this payoff transformation is applied to the Centipede Game, the entire extensive-form tree undergoes a radical transformation. If Player 2 possesses a sufficiently high guilt parameter β2, the personal psychological utility derived from defecting at Node M and leaving Player 1 with an inferior payoff is strictly lower than the utility derived from passing and sharing an equal, expanded cooperative payoff. Consequently, for an inequity-averse Player 2, passing at the terminal node is no longer an irrational sacrifice—it is an optimal, utility-maximizing act! Furthermore, if Player 1 assigns a sufficiently high subjective probability to Player 2 possessing these other-regarding preferences, Player 1’s rational choice at Node M − 1 shifts from defection to cooperation. By replacing cold monetary payoffs with psychologically authentic utility functions, behavioral economists proved that deep cooperation in the Centipede Game can be entirely compatible with rational optimization.

7.2 Trust, Reciprocity, and Psychological Game Theory

While distributional models like Fehr-Schmidt capture static preferences for fairness, they fail to model the dynamic, emotional mechanisms of reciprocity. In human social interaction, people do not merely evaluate the final allocation of dollars; they care deeply about the intentions behind the actions that produced that allocation. This reality necessitated the development of Psychological Game Theory, pioneered by Matthew Rabin (1993) and extended to extensive-form games by Martin Dufwenberg and Georg Kirchsteiger (2004).

In sequential interactions, passing the move in the Centipede Game is not a passive event; it is an active, highly visible demonstration of trust. At Node 1, Player 1 holds the unilateral power to take the money and ensure that Player 2 receives nothing. By choosing to pass, Player 1 voluntarily foregoes an immediate, guaranteed payoff and exposes themselves to significant vulnerability, placing their financial welfare directly into Player 2’s hands. This act conveys an unambiguous, benevolent intention: it is an expensive gift of trust designed to expand the pie for both players.

Under psychological game theory, human beings are hardwired with a universal norm of positive reciprocity: an innate desire to reward kindness with kindness, and to punish hostility with hostility. When Player 2 receives the move at Node 2, they perceive Player 1’s pass as an act of genuine generosity. This psychological perception triggers a powerful reciprocal obligation. To exploit this gift by immediately defecting at Node 2 would induce acute internal guilt and violate deeply internalized social contracts. Passing at Node 2 is thus fueled by sequential reciprocity: an ongoing, dynamic dialogue of mutual trust, where each successive pass elevates the psychological debt of kindness, driving the players deeper and deeper into the game tree.

7.3 Reputation and Incomplete Information Models (The Kreps-Milgrom-Roberts-Wilson Framework)

An alternative explanation of the Centipede paradox was pioneered by David Kreps, Paul Milgrom, John Roberts, and Robert Wilson (1982) through their celebrated KMRW Gang of Four framework. The profound genius of the KMRW insight is that it does not require all, or even most, players to be genuinely altruistic. It demonstrates that the mere existence of a tiny, infinitesimal probability of an altruistic “type” within the population can compel completely cold, selfish, hyper-rational actors to mimic cooperative behavior until the closing moments of the game.

The mathematical architecture operates through Bayesian game theory with incomplete information:

  • Suppose that the vast majority of the population consists of standard, hyper-rational, selfish utility maximizers (Type R).
  • However, there exists a small subjective probability, ε > 0 (say, 5%), that any randomly encountered player is an intrinsically cooperative “altruistic type” (Type A) who will unconditionally pass at every node.

Now, consider the strategic calculations of a completely selfish, rational Player 1 at Node 1. If Player 1 defects immediately, they receive pennies. What happens if Player 1 passes? By passing, Player 1 signals to Player 2 that there is a possibility they might be an altruistic Type A. Player 2, upon receiving the move at Node 2, updates their beliefs via Bayes’ rule. Even if Player 2 is purely selfish, Player 2 realizes that if they pass at Node 2, they can keep the game going and harvest a massively larger payoff downstream. Therefore, a rational, selfish Player 2 has a powerful financial incentive to pass in order to encourage Player 1 to pass again!

This creates a lucrative reputation-building equilibrium. Selfish players actively masquerade as cooperative types. They wear the mask of altruism not out of moral goodness, but because pretending to be cooperative generates vast financial surpluses. This cooperative facade holds firm across the majority of the game tree. However, as the terminal node approaches, the expected future gains from sustaining the illusion diminish relative to the immediate temptation of defection. At that threshold, the Bayesian equilibrium unravels, and players race to defect. The KMRW framework matches McKelvey and Palfrey’s empirical data: sustained cooperation through the early and intermediate nodes, followed by a sharp outbreak of defections near the end of the tree.

8. Stochastic Choice and Quantal Response Equilibrium (QRE) in the Centipede Game

8.1 Principles of Quantal Response and Bounded Rationality Models

Despite the insights provided by social preference and reputation models, an alternative explanation emerged directly from the mathematics of decision errors. Classical game theory assumes that players execute deterministic best responses: if action A yields an expected payoff of $10.01 and action B yields $10.00, the probability that the player chooses A is precisely 1.0, and the probability of choosing B is 0.0. This razor-sharp sensitivity to microscopic payoff differentials is fundamentally at odds with empirical human psychology, where decision-making is intrinsically stochastic, noisy, and vulnerable to cognitive trembles.

To capture this reality, Richard McKelvey and Thomas Palfrey (1995, 1998) introduced the revolutionary concept of the Quantal Response Equilibrium (QRE), and its sequential counterpart, the Agent Quantal Response Equilibrium (AQRE). In QRE, the assumption of deterministic optimization is replaced by a stochastic choice rule, typically formalized via a logit choice function. At any decision node, the probability of selecting a given action is a smooth, monotonically increasing function of that action’s expected utility.

For an agent facing a choice between actions with expected payoffs Ui, the probability P(aj) of selecting action aj is defined by the classic logit formula:

P(aj) = exp(λ U(aj)) / ∑k exp(λ U(ak))

The critical parameter in this formulation is λ (lambda), representing the precision of choice or the degree of rationality:

  • As λ → 0, the precision collapses to zero. Payoff differentials become irrelevant, and the player behaves with complete, uniform randomness (50/50 probability across binary options).
  • As λ → ∞, the noise vanishes entirely, and the choice rule converges asymptotically to the hyper-rational, deterministic best-response of classical game theory.
  • At intermediate, empirical values (0 < λ < ∞), players exhibit better-response behavior: they are statistically more likely to choose actions that yield higher expected payoffs, but they occasionally make errors, with the probability of an error being strictly inverse to the cost of making it.

8.2 Structural Estimation on McKelvey and Palfrey’s Empirical Datasets

When McKelvey and Palfrey applied Agent Quantal Response Equilibrium (AQRE) to their empirical Centipede Game data, the mathematical fit was nothing short of miraculous. The model resolved the backward induction paradox without requiring any exotic assumptions about altruism, fairness, or social preferences. It explained the observed behavior purely as a consequence of rational actors operating in an environment containing a predictable background rate of decision noise.

To understand the profound structural mechanics of AQRE in the Centipede Game, one must trace how stochastic errors propagate backward through the extensive-form tree:

  • Consider the final decision node, Node M. Even if Player 2 is purely selfish and prefers to Take, under logit choice with a finite precision parameter λ, there is a small, non-zero probability εM that Player 2 will accidentally or erroneously choose to Pass.
  • Now, examine the perspective of Player 1 at Node M − 1. In standard backward induction, Player 1 assumes that Player 2 will pass with a probability of strictly zero. Under AQRE, Player 1 knows that Player 2 will pass with probability εM > 0. If the cooperative payoff resulting from a pass is sufficiently massive relative to the small cost of being defected upon, the expected value of passing at Node M − 1 becomes strictly higher than the value of defecting!
  • Consequently, a rational, self-interested Player 1 will choose to Pass at Node M − 1 with high probability!
  • Continuing this backward calculus to Node M − 2, Player 2 now faces an upstream environment where Player 1 is highly likely to pass. This increases the expected value of passing for Player 2 at Node M − 2.

The noise at the final node acts as a behavioral foundation that supports deep, robust towers of cooperation throughout the entire upstream tree. By fitting a single structural precision parameter λ across the experimental datasets, McKelvey and Palfrey demonstrated that AQRE could predict the exact node-by-node termination frequencies with extraordinary statistical accuracy. The unraveling of backward induction was revealed to be a direct mathematical consequence of cumulative risk aggregation across sequential human nodes.

8.3 Limitations and Critiques of Stochastic Decision Frameworks

Despite its mathematical elegance and immense descriptive success, the Quantal Response Equilibrium framework has faced significant theoretical and methodological criticisms within the broader behavioral economics community.

The foremost critique centers on the problem of post-hoc curve fitting and parameter endogeneity. Critics, including Colin Camerer and Vincent Crawford, have pointed out that because the precision parameter λ is typically estimated econometrically after the laboratory data has already been collected, the model risks becoming an exercise in sophisticated curve-fitting. Because λ encapsulates all forms of unmodeled variance—ranging from perceptual mistakes and computational lapses to unobserved altruistic motives and risk preferences—it can act as a theoretical sponge that absorbs any systematic discrepancy between theory and data without providing an explicit causal mechanism.

Furthermore, AQRE frequently struggles with out-of-sample predictive generalizability. A value of λ estimated with exquisite precision in a four-node centipede game often fails to accurately predict choice distributions when applied to a ten-node game, or when transplanted into an entirely different game topology, such as a signaling game or a public goods environment. Finally, pure stochastic frameworks present a profound philosophical ambiguity: they treat human cooperation essentially as an “error.” To classify a deeply cooperative move that generates vast mutual social wealth as a cognitive mistake—on par with misreading an interface or clicking the wrong button—misses the rich psychological, evolutionary, and normative dimensions of human prosociality.

9. Comparative Analysis: Rosemarie Nagel’s Guessing Game vs. Richard’s Centipede Game

9.1 Strategic Form (Simultaneous) Versus Extensive Form (Sequential) Cognition

A rigorous comparative analysis of Rosemarie Nagel’s Guessing Game and Richard Rosenthal’s Centipede Game illuminates the fundamental cognitive dichotomy that separates normal-form and extensive-form strategic interactions. Though both games are designed to expose the empirical breakdown of infinite deductive reasoning, they engage distinct psychological architectures, impose disparate forms of cognitive load, and process feedback through radically divergent channels.

Strategic Dimension Rosemarie Nagel’s Guessing Game Rosenthal / McKelvey-Palfrey Centipede Game
Structural Form Normal (Strategic) Form; Simultaneous Play Extensive Form; Dynamic Sequential Play
Equilibrium Mechanism Iterated Elimination of Dominated Strategies (IEDS) Backward Induction / Subgame Perfection (SPNE)
Cognitive Architecture Parallel recursive simulation; population belief modeling Temporal backward projection; tree folding; trust calibration
Information Environment Symmetric ignorance; choices made in complete isolation Perfect information; explicit observable historical decisions
Primary Behavioral Friction Bounded depth of reasoning (Level-k exhaustion) Social preferences, intentional trust, error propagation (QRE)
Vulnerability to Others Collective aggregation; insulated by the law of large numbers Pairwise vulnerability; total exposure to partner’s defection

In Nagel’s Guessing Game, the decision-maker operates under conditions of symmetric parallel uncertainty. The individual must engage in an internal, computational simulation of an entire population. The cognitive challenge is static and mathematical: How many layers of “they think that I think” can I compute before my working memory halts? Crucially, an individual in the guessing game is shielded by the aggregating mechanics of the arithmetic mean; the idiosyncratic madness or genius of any single opponent is diluted across the entire cohort.

In the Centipede Game, the cognitive landscape is entirely different. Reasoning is not simultaneous, but deeply temporal and sequential. A player does not simulate a faceless population average; they face a specific, living human being across an expanding timeline. The cognitive burden centers on dynamic intentional attribution. Every historical action by the counterpart sends a profound, visible signal. Passing at Node 2 is not a data point in an aggregate statistic; it is a direct personal gift of trust that forces the recipient to confront complex social norms of reciprocity, fear of exploitation, and strategic betrayal. The Centipede Game converts strategic thinking into a high-stakes psychological drama unfolding one step at a time.

9.2 Iterated Elimination of Dominated Strategies Versus Backward Induction

From an abstract, purely mathematical standpoint, the Iterated Elimination of Dominated Strategies (IEDS) executed in Nagel’s game and the Backward Induction executed in Rosenthal’s game are theoretical cousins. Both represent recursive deductive algorithms that systematically prune away suboptimal strategies under the strict assumption of common knowledge of rationality. In a finite world of classical logic, both algorithms inexorably drive the theoretical solution to a single, radical boundary: zero in the guessing game, and immediate first-node defection in the centipede game.

Yet, when subjected to human empirical cognition, these two analytical mechanics behave in fundamentally different ways:

  • In Nagel’s game, the iterations of IEDS occur entirely in the abstract space of hypothetical thoughts. To discard strategies above 66.67, a player does not need to observe anyone actually playing a strategy; the deduction is purely prospective. Because the game is simultaneous, there are no actual, observable “rounds” within a single trial. A player stops at Level-1 or Level-2 simply because their cognitive engine lacks the working memory or inclination to calculate further.
  • In Rosenthal’s game, backward induction requires a player to mentally project to a terminal node that may never be physically reached. The player must calculate what would happen under counterfactual conditions that their own immediate choice might actively prevent from occurring!

This creates a massive empirical divergence. While bounded reasoning in the guessing game manifests as a gentle, graceful cognitive stopping point (yielding stable choices around 33 or 22), bounded reasoning in the centipede game directly generates an epistemic crisis. Reaching Node 2 directly shatters the premise of backward induction in real time, forcing the players out of the deductive domain of classical game theory and into the inductive domain of psychological signaling, trust, and evolutionary heuristics.

9.3 The Intersection of Beliefs, Information Structures, and Coordination

When synthesizing the overarching lessons of Nagel’s and Rosenthal’s experimental breakthroughs, we uncover fundamental truths regarding how human beings formulate beliefs, navigate information structures, and achieve social coordination under radical strategic uncertainty.

Both games conclusively demonstrate that rationality is profoundly contextual and relational. In both environments, an agent who acts in strict accordance with classical equilibrium theory—submitting 0 in a one-shot guessing game, or defecting at Node 1 in the centipede game—performs terribly in practice. In the guessing game, choosing 0 guarantees that the player will lose the prize to a Level-2 thinker choosing 22. In the centipede game, defecting at Node 1 guarantees that the player walks away with pennies, forfeiting hundreds of dollars in collective wealth successfully unlocked by conditionally cooperative pairs. Classical game theory defines rationality in a vacuum of self-referential logic; empirical success requires what game theorists now call ecological rationality—the capacity to accurately model, anticipate, and harmonize with the bounded, psychological realities of the real human beings who constitute the strategic environment.

Moreover, both paradigms expose the decisive stabilizing power of common information and dynamic feedback. In Nagel’s game, public feedback regarding the aggregate mean rapidly aligns subjective beliefs, creating a unified directional vector that drives heterogeneous Level-k thinkers toward the Nash equilibrium over successive periods. In the centipede game, the physical passage of the move serves as a powerful coordinating device, creating an emergent track record of mutual benevolence that temporarily suspends the toxic unraveling of backward induction. Together, the works of Rosemarie Nagel and Richard Rosenthal establish that strategic coordination in human civilization is not born of cold, transfinite deduction, but is meticulously engineered through shared social norms, dynamic feedback loops, and the resilient human capacity to navigate bounded rationality.

10. Methodological Rigor and Laboratory Protocols in Behavioral Economics

10.1 Experimental Controls, Anonymity, and Financial Incentives

The profound discoveries generated by Nagel and Rosenthal’s experimental paradigms were made possible only through the meticulous development of experimental methodology in economics. Unlike traditional psychological experiments, which often rely on hypothetical scenarios, survey instruments, or deceptive scripts, experimental economics established an unyielding commitment to methodological transparency, salient financial incentives, and rigorous control.

At the center of this empirical philosophy stands Vernon Smith’s Induced Value Theory. To ensure that an experiment measures authentic strategic reasoning rather than casual, unmotivated guessing, the experimenter must establish control over the subjects’ internal preferences. This is achieved by utilizing tangible, salient financial rewards that dominate the subjective, psychological transaction costs of participating. In both Nagel’s beauty contest and McKelvey and Palfrey’s centipede experiments, every decision carried immediate financial consequences. Players were paid substantial cash rewards directly proportional to their strategic performance. This incentive structure ensures that observed deviations from equilibrium are not the product of boredom or indifference, but represent genuine cognitive boundaries operating under optimal motivation.

Equally critical is the universal prohibition against deception. In modern experimental economics, researchers are ethically and methodologically forbidden from misleading subjects about game rules, counterpart identities, or payoff calculations. If an experimental protocol states that a player is randomly matched with an anonymous peer, that matching must be executed with mathematical fidelity. This ironclad integrity ensures that participants form subjective beliefs based purely on the structural rules of the game, rather than suspecting that the experimenter is secretly manipulating the environment. Combined with double-blind protocols—where neither the other players nor the experimenter can link an individual’s identity to their specific choices—these controls create an environment where strategic reasoning can be isolated and measured with surgical precision.

10.2 Replication, Robustness, and Cross-Cultural Experimental Demographics

A defining hallmark of both the Guessing Game and the Centipede Game is their extraordinary empirical robustness. In an era where social sciences have been rocked by systemic replication crises, the findings of Nagel and McKelvey-Palfrey have stood as unassailable pillars of empirical reliability, successfully replicated across hundreds of independent laboratories, academic disciplines, and cultural contexts worldwide.

The p-beauty contest has been administered across an extraordinary demographic spectrum:

  • Undergraduate university students from elite institutions (Caltech, Universitat Pompeu Fabra, Oxford).
  • High-powered financial professionals, including Wall Street portfolio managers, foreign exchange traders, and central bankers.
  • World-renowned research mathematicians, computer scientists, and professional game theorists.
  • General populations via nationwide newspaper competitions involving tens of thousands of citizens in Spain, Germany, and the United Kingdom.

The results across these vastly disparate cohorts are remarkably consistent. While professional traders and trained game theorists display a slightly higher average depth of reasoning—yielding higher frequencies of Level-2 and Level-3 choices—the modal clusters at 33 and 22 persist stubbornly across all demographics. Even cohorts composed entirely of professional economists fail to converge to zero in first-period play! The mean depth of reasoning remains bounded between 1.5 and 2.5 steps across the globe, proving that cognitive hierarchy is a fundamental cognitive universal of the human brain rather than a byproduct of education, wealth, or cultural background.

Similarly, the Centipede Game has undergone exhaustive cross-cultural and demographic testing. Experiments conducted across North America, Western Europe, and East Asia demonstrate the universal endurance of initial cooperation. While regional variations emerge in the exact node of peak defection—reflecting subtle cultural differences in generalized social trust and institutional norms—the fundamental phenomenon remains invariant: immediate first-node defection is universally rejected, and backward induction fails to describe sequential human play across all tested societies.

10.3 Design Innovations: Constant Sum, Constant Ratio, and Infinite Horizon Extensions

To deepen the theoretical insights yielded by the baseline experiments, researchers continuously innovate upon the physical and mathematical architecture of the games. In the extensive-form domain, theorists developed critical variations of Rosenthal’s centipede framework to isolate the precise mathematical drivers of human cooperation:

  • Constant-Sum Centipede Games: In these implementations, passing does not expand the financial pot; the total wealth remains strictly fixed throughout the entire tree. These experiments proved that when the cooperative surplus is removed, players defect significantly earlier, confirming that human cooperation in the baseline game is powerfully fueled by a rational desire to harvest the massive gains from pot expansion.
  • Constant-Ratio Centipede Games: Payoffs grow exponentially at each node by a constant multiplier (e.g., doubling at every step), creating an environment where the absolute opportunity cost of defecting explodes as the game progresses. These designs revealed that players will sustain cooperation across vast numbers of nodes when the financial return on mutual trust is high.
  • Probabilistic Termination Rules: To bridge the gap between finite and infinite-horizon games, researchers implemented protocols where the game does not conclude at an arbitrary, fixed terminal node. Instead, after each pass, a randomizing device (such as a computerized dice roll) determines whether the game continues to the next node or terminates exogenously with a fixed probability (e.g., δ = 0.95). This design mirrors the classical discounted infinite-horizon model, successfully eliminating the artificial, backward-unraveling boundary conditions imposed by fixed terminal nodes.

Simultaneously, within the guessing game literature, researchers developed linear and non-linear asymmetric scoring rules, where participants are penalized differently for guessing above versus below the target. Others implemented real-time continuous-clock environments, where players can observe the shifting collective mean on a digital dashboard and continuously modify their entries before a randomized countdown timer freezes the market. These technological and structural innovations continue to expand the empirical frontier, testing the resilience of bounded rationality across increasingly complex institutional landscapes.

11. Macroeconomic, Financial, and Strategic Policy Implications

11.1 Asset Price Bubbles and Herd Mentality in Financial Markets

The direct mapping between Rosemarie Nagel’s p-beauty contest and real-world capital markets provides macroeconomists and financial analysts with a vital framework for understanding the mechanics of asset price bubbles, speculative manias, and systemic market panics. Classical finance theory, grounded in the Efficient Market Hypothesis (EMH), asserts that the price of an asset continuously reflects its fundamental discounted cash flows. Under this neoclassical view, bubbles should be impossible, as rational arbitrageurs will instantly short-sell any overvalued asset, driving its valuation back to intrinsic equilibrium.

However, real financial markets operate as gigantic, highly leveraged Keynesian beauty contests. A professional trader’s performance is rarely evaluated against an asset’s discounted cash flows thirty years into the future. Instead, fund managers are evaluated against their quarterly benchmarks and their peer performance. In this environment, the central question is not what an asset is truly worth, but what other market participants will be willing to pay for it next week, next month, or next quarter.

This dynamic produces what behavioral finance classifies as the “Riding the Bubble” phenomenon, a direct real-world manifestation of Level-k reasoning:

  • A Level-1 investor observes an asset (such as dot-com equities in 1999, residential subprime derivatives in 2006, or speculative cryptocurrencies) and buys it simply because its price is going up.
  • A sophisticated Level-2 or Level-3 institutional investor fully realizes that the asset is wildly overvalued relative to its fundamental metrics. Yet, this sophisticated investor does not short-sell the asset! They know that the market is flooded with Level-1 capital that will continue driving the price upward.
  • Consequently, the rational, utility-maximizing move for the sophisticated investor is to buy the bubble, ride the upward price trajectory, and plan to exit the market one step before the aggregate herd decides to liquidate.

The systemic catastrophe occurs because calculating that precise exit step requires common knowledge of the population’s cognitive hierarchy, which does not exist. Every trader believes that they possess higher strategic sophistication (a higher k) than the average market participant, convincing themselves that they will successfully dump their holdings right at the market’s peak. When the aggregate valuation inevitably strikes an exogenous arithmetic ceiling, the backward-unraveling logic observed in multi-period guessing games triggers instantaneously. The market experiences a non-linear, catastrophic collapse: an avalanche of simultaneous sell orders as all cognitive levels rush for the exit simultaneously, freezing market liquidity and sparking systemic macroeconomic crises.

11.2 Strategic Brinkmanship, Geopolitical Escalation, and Nuclear Deterrence

While the guessing game illuminates financial markets, Richard Rosenthal’s Centipede Game serves as a foundational structural metaphor for the high-stakes arena of geopolitical brinkmanship, diplomatic bargaining, and military escalation. In international relations, conflicts rarely manifest as sudden, unannounced total wars. Instead, crises unfold as alternating, sequential extensive-form games of mutual testing, signaling, and calculated escalation.

Consider a territorial or diplomatic standoff between two nuclear-armed superpowers (such as the United States and the Soviet Union during the Cuban Missile Crisis, or contemporary maritime tensions in the South China Sea). The strategic landscape closely mirrors a high-stakes Centipede Game:

  • At each juncture, a nation can choose to De-escalate/Pass, maintaining diplomatic dialogue, preserving mutual peace, and passing the strategic initiative to the adversary.
  • Alternatively, the state can choose to Escalate/Defect, seizing a tactical military advantage, deploying naval blockades, or striking an adversary’s outpost before the adversary can act.

If geopolitical leaders were to operate according to the strict, cold mathematics of Subgame Perfect Nash Equilibrium, modern civilization would be completely unsustainable. Backward induction dictates that because an adversary will rationally seek to dominate the ultimate terminal stage of a crisis, a state must preemptively execute a devastating, unprovoked surprise strike at the very first possible moment. Subgame perfection in international relations leads straight to preemptive annihilation.

The survival of the geopolitical order relies upon the exact same behavioral mechanisms that preserve cooperation in McKelvey and Palfrey’s laboratory: conditional trust, intentional signaling, and sequential reciprocity. When one nation deliberately refrains from exploiting an adversary’s temporary vulnerability, that calculated pass signals benevolence and a commitment to stability. It establishes an implicit bilateral contract that enables both nations to navigate deep into the crisis tree without triggering catastrophic defection. However, the Centipede Game also sounds a terrifying warning for geopolitical architects: because empirical cooperation tends to break down near the terminal nodes of an extensive interaction, any geopolitical architecture that permits an obvious “final step”—such as a rigid, expiring treaty deadline or a zero-sum territorial boundary—risks triggering preemptive backward unraveling, culminating in rapid, catastrophic military escalation.

11.3 Institutional Design, Mechanism Design, and Regulatory Architecture

The transformative insights of experimental game theory have revolutionized modern Mechanism Design—the branch of economic science dedicated to engineering optimal rules, laws, and market institutions. Classical mechanism design, founded by Leonid Hurwicz, Eric Maskin, and Roger Myerson, traditionally engineered auctions, matching algorithms, and regulatory frameworks under the ironclad assumption that all participating economic agents would play the unique Subgame Perfect Nash Equilibrium of the designed mechanism.

The empirical revelations of Nagel and Rosenthal proved that classical mechanism design was built on fragile foundations. A mechanism that is mathematically optimal under the assumption of infinite iterative rationality can fail catastrophically when populated by real human beings operating under Level-1 or Level-2 cognitive constraints:

  • In spectrum and government asset auctions, mechanisms that rely on complex chains of dynamic backward induction routinely collapse into disastrous outcomes, plagued by widespread bidder confusion, unexpected default rates, and massive revenue shortfalls.
  • Consequently, modern institutional designers prioritize the engineering of strategy-proof mechanisms that are robust to bounded rationality. Prominent among these is the concept of Obviously Strategy-Proof (OSP) mechanisms, pioneered by Shengwu Li, where an optimal choice can be recognized through immediate, simple dominance without requiring any multi-step backward induction or complex simulation of others’ mental states.

In central banking and monetary policy, the recognition of Level-k dynamics has fundamentally transformed modern communications and forward guidance. Central banks no longer assume that announcing an interest rate trajectory will automatically trigger hyper-rational, instantaneous rational-expectations adjustments across the entire macroeconomy. Instead, monetary policymakers explicitly recognize that market participants possess heterogeneous reasoning levels (τ ≈ 1.5). Forward guidance is now deliberately engineered as a powerful, focal narrative device designed to guide and anchor the beliefs of Level-1 and Level-2 market actors, actively preventing the speculative divergence and coordinating economic expectations around macroeconomic stability.

12. Epistemological Synthesis and the Future of Experimental Strategic Analysis

12.1 Re-evaluating the Axioms of Rationality in Social Sciences

The intellectual trajectories initiated by Rosemarie Nagel’s Guessing Game and Richard Rosenthal’s Centipede Game have forced a profound epistemological re-evaluation across the social sciences. For over half a century, the axiomatic framework of homo economicus served as the undisputed core of theoretical economics, celebrated for its internal consistency and mathematical elegance. Yet, the persistent empirical realities uncovered in the laboratory have irrevocably challenged the foundational dogma that pure deductive rationality describes the human condition.

This empirical revolution has not rendered classical game theory obsolete. Rather, it has fundamentally clarified its scientific purpose. The mathematical benchmarks of Nash equilibrium, backward induction, and subgame perfection do not represent descriptive portraits of empirical reality; they serve as indispensable normative baselines. Just as the idealized frictionless planes and perfect vacuums of Newtonian physics provide the essential mathematical reference points against which real-world air resistance and friction are measured, the equilibria of classical game theory define the pristine coordinates against which the frictions, bounds, and social preferences of human psychology can be rigorously mapped and quantified.

Economics has thus evolved from a closed, deductive mathematical doctrine into an open, empirical behavioral science. The modern paradigm integrates the formal analytical scaffolding of mathematics with the rich descriptive insights of cognitive psychology, evolutionary anthropology, and neurobiology. Bounded rationality is no longer dismissed as an embarrassing collection of human mistakes or anomalies; it is recognized as the structural, evolutionary logic of human cognition—an architecture optimized through millions of years of natural selection to solve complex, uncertain social challenges using fast, frugal, and remarkably robust heuristics.

12.2 Artificial Intelligence, Algorithmic Agents, and Equilibrium Selection

As the global economy transitions into an era dominated by autonomous algorithmic systems, high-frequency machine trading, and artificial intelligence, the experimental paradigms of Nagel and Rosenthal are finding urgent new applications at the intersection of computer science and economics. Modern markets are no longer populated exclusively by human minds; they are increasingly governed by deep reinforcement learning agents, neural networks, and automated algorithmic decision-makers.

Recent experiments pitting cutting-edge artificial intelligence agents against human subjects in the Guessing Game and Centipede Game have revealed fascinating strategic dynamics:

  • When modern reinforcement learning algorithms (such as deep Q-networks or actor-critic architectures) are trained through self-play in the p-beauty contest, they rapidly learn to execute the complete chain of iterated elimination of dominated strategies, converging flawlessly to the Nash equilibrium of 0.
  • However, when these mathematically “flawless” AI agents are placed into hybrid experimental environments populated by actual human beings, the algorithmic agents consistently lose! By choosing 0, the AI executes a hyper-rational calculation that is fundamentally mismatched to the empirical Level-1 and Level-2 human population.

To survive and dominate in real-world human environments, artificial intelligence architectures must be explicitly engineered with Theory of Mind modules. Contemporary computer scientists are integrating structural Cognitive Hierarchy and Quantal Response Equilibrium models directly into machine learning pipelines. By deploying algorithms capable of dynamically estimating the specific τ (tau) or λ (lambda) of their human counterparts through real-time Bayesian updating, machine agents can strategically calculate optimal best-responses to human bounded rationality—exploiting the predictable, structured bounds of human thinking in asset markets, negotiations, and strategic auctions.

12.3 Open Questions and Future Frontiers in Behavioral Game Theory

Despite the immense theoretical and empirical progress achieved over the past four decades, the intellectual frontiers opened by Rosemarie Nagel and Richard Rosenthal remain vibrant with unanswered questions, deep methodological debates, and exciting paths for future research. Behavioral game theorists continue to grapple with fundamental foundational challenges that resist easy resolution.

One of the most pressing open frontiers is the development of a Unified Theory of Strategic Reasoning capable of seamlessly synthesizing static and dynamic environments into a single, comprehensive mathematical framework. Currently, the discipline deploys Level-k and Cognitive Hierarchy models to explain normal-form simultaneous games, while relying on Quantal Response Equilibrium, Bayesian reputation models, and social preference frameworks to explain dynamic extensive-form games. The development of a single, coherent cognitive architecture—one that simultaneously captures bounded recursive depth, stochastic execution errors, dynamic forward belief updating, and visceral other-regarding preferences without sacrificing mathematical parsimony—remains the holy grail of modern behavioral economics.

Simultaneously, the integration of ultra-high-resolution, continuous-time neuro-computational tracking promises to unlock the biological mechanics of decision-making. Future laboratories will track intracranial neural oscillations, pupillometry, and galvanic skin responses in real time as human participants navigate high-stakes, multi-node strategic trees. These physiological frontiers will allow researchers to observe the precise millisecond when working memory halts, when social empathy overrides financial greed, and when the instinct to cooperate gives way to the preemptive fear of betrayal.

Ultimately, the enduring legacies of Rosemarie Nagel and Richard Rosenthal lie in their transformative impact on our understanding of ourselves. By demonstrating that the human mind is neither a broken computer nor an omniscient calculating machine, their pioneering experiments illuminated the profound, beautiful complexity of human sociality. They showed that our bounded rationality is not an intellectual curse, but the very mechanism through which humanity navigates uncertainty, transcends cold theoretical boundaries, and weaves the enduring fabric of social trust and civilizational cooperation.

Conclusion

The trajectory of behavioral game theory—sparked by the conceptual formulations of the p-beauty contest and the Centipede Game—has irrevocably altered the landscape of economic science. Rosemarie Nagel provided the discipline with an empirical mirror, demonstrating that when human beings are asked to anticipate the thoughts of their peers, they do not engage in infinite mathematical recursion; instead, they step gracefully and predictably through one or two discrete layers of reasoning. Richard Rosenthal presented game theory with an empirical paradox, illustrating that the cold, destructive logic of backward induction is consistently defeated by human prosociality, trust, and the stabilizing realities of stochastic decision noise.

Together, these paradigms bridge the divide between theoretical mathematics and the empirical realities of human psychology. They prove that human strategic behavior is not an undisciplined wilderness of chaotic errors, but a deeply structured, lawful realm characterized by finite cognitive depths, rich social preferences, and adaptive learning dynamics. As social scientists, mechanism designers, and technological architects confront the monumental challenges of the twenty-first century—from the regulation of algorithmic markets to the fragile dynamics of international geopolitics—the insights forged by Nagel, Rosenthal, McKelvey, and Palfrey stand as indispensable beacons. They remind us that any model of society that fails to account for the real, bounded nature of human cognition is doomed to failure, and that the ultimate promise of economic science lies in understanding humanity as it truly is.

References

Rate This Content

0.0 / 5 0 votes

Cite This Article

memjavad (2026, September 12). Guessing Game – Rosemarie Nagel The Centipede Game Experiment – Richard. PSYCHOLOGICAL DATABASE. https://en.arabpsychology.com/experiments/guessing-game-rosemarie-nagel-centipede-game-experiment-richard/
memjavad. “Guessing Game – Rosemarie Nagel The Centipede Game Experiment – Richard.” PSYCHOLOGICAL DATABASE, 12 September 2026, https://en.arabpsychology.com/experiments/guessing-game-rosemarie-nagel-centipede-game-experiment-richard/.
memjavad. “Guessing Game – Rosemarie Nagel The Centipede Game Experiment – Richard.” PSYCHOLOGICAL DATABASE. September 12, 2026. https://en.arabpsychology.com/experiments/guessing-game-rosemarie-nagel-centipede-game-experiment-richard/.