For more than half a century, non-cooperative game theory occupied a curious epistemological position within economics. Formalized through the pioneering work of John Nash, John von Neumann, and Oskar Morgenstern, the discipline was built upon an uncompromising architecture of hyper-rationality. In this theoretical landscape, agents were assumed not only to possess unbounded computational capacity, but also to maintain mutually consistent, rational expectations regarding one another’s actions. The standard analytical benchmark—the Nash equilibrium—presupposed that every participant perfectly anticipates their opponents’ optimal strategies and selects a mutual best response. When laboratory experiments began to systematically test these axioms in the late twentieth century, however, a profound dissonance emerged. Human subjects frequently and systematically failed to play equilibrium strategies, particularly in initial, unlearned encounters. Instead, their choices revealed subtle, structured patterns of cognition that defied textbook equilibrium calculations, exposing the inadequacy of treating game-theoretic rationality as an unassailable empirical fact.
Faced with this empirical rift, mainstream economic methodology long retreated behind the classical revealed-preference paradigm and Milton Friedman’s instrumentalist “as-if” doctrine. According to this view, the internal, algorithmic mechanics of human reasoning were considered an unobservable “black box,” irrelevant so long as aggregate market outcomes approximated equilibrium states. Yet in strategic contexts, revealed preference suffers from a fundamental identification problem: vastly different cognitive heuristics, differing depth-of-reasoning levels, and genuine equilibrium calculations frequently select the exact same observable actions across standard normal-form matrices. Without observing the procedural, intermediate steps of cognition—how agents acquire, prioritize, and process information—economists remained trapped in post-hoc curve fitting, unable to discern whether an equilibrium choice was the result of sophisticated fixed-point reasoning or a sheer coincidence born of simple, myopic heuristics.
It was against this intellectual backdrop that Miguel Costa-Gomes, Vincent P. Crawford, and Bruno Broseta published their monumental 2001 Econometrica paper, “Cognition and Behavior in Normal-Form Games: An Experimental Study.” In an ambitious theoretical and empirical undertaking, the authors synthesized structural econometric modeling, non-equilibrium decision theory, and high-resolution process-tracing technology. By tracking subjects’ sequential eye movements and information acquisition across eighteen carefully calibrated normal-form games using the MouseLab computer interface, Costa-Gomes, Crawford, and Broseta dismantled the neoclassical black box. Their inquiry established a rigorous empirical taxonomy of strategic decision types—ranging from naive heuristics to iterated dominance and bounded Level-k hierarchies—and inaugurated the modern era of “thinking experiments,” transforming game theory from a purely normative mathematical exercise into a deeply empirical cognitive science.
1. Foundations of Experimental Game Theory and the Cognition Problem
1.1 The Neoclassical Assumption of Hyper-Rationality in Strategic Settings
The neoclassical foundation of non-cooperative game theory rests upon an interlocking set of epistemic and behavioral axioms. At its core lies the assumption of hyper-rationality: each player is assumed to possess a complete, transitive preference ordering over outcomes, unbounded mathematical capacity to calculate expected payoffs across arbitrary probability distributions, and the cognitive facility to execute infinite chains of counterfactual reasoning. These individual capacities are elevated to the collective level via the assumption of common knowledge of rationality (CKR), which asserts that all players are rational, all players know that all players are rational, and this mutual awareness continues ad infinitum. Under CKR, strategic interaction collapses to the Nash equilibrium concept, where beliefs are mutually consistent, self-fulfilling, and invariant to strategic surprises.
Despite the elegance of these axioms, the historical development of experimental economics systematically revealed persistent, non-random deviations from Nash equilibrium predictions. When placed in laboratory environments, intelligent human subjects consistently violate the behavioral prescriptions of backward induction, iterated deletion of strictly dominated strategies, and mixed-strategy equilibrium. Landmark investigations in the 1980s and 1990s demonstrated that players often choose strictly dominated actions, fail to coordinate on Pareto-dominant equilibria in stag hunt and coordination games, and exhibit non-negligible rates of out-of-equilibrium play in finite centipede and bargaining games.
These empirical anomalies underscored the fundamental limitation of backward induction and iterated dominance as descriptive behavioral models. While mathematically robust, backward induction assumes that agents can execute cognitive cascades through extensive-form game trees without facing computational frictions, working-memory constraints, or doubts regarding their opponents’ structural rationality. In reality, human cognition is intrinsically bounded. As Herbert A. Simon emphasized in his seminal formulations of bounded rationality, decision-makers confronted with complex, information-rich strategic environments do not optimize globally. Instead, they rely on procedural heuristics, satisficing search algorithms, and truncated cognitive steps. Integrating bounded rationality into non-cooperative game theory required abandoning the assumption of instantaneous, frictionless equilibrium coordination in favor of explicit, micro-founded models of human cognition.
1.2 The Black Box of Strategic Choice and the Need for Process Data
For decades, empirical economics remained anchored in the revealed-preference framework formalized by Paul Samuelson. In the context of strategic matrix games, this methodology dictated that researchers observe only the final action selected by a subject—such as selecting Row A over Row B. Economists would subsequently infer the underlying preference ordering or belief structure based entirely on this terminal choice. However, in non-cooperative games, this approach encounters an insurmountable identification problem. A single strategic move can frequently be justified by multiple, theoretically distinct decision rules. For example, a player who chooses an action corresponding to the unique Nash equilibrium might have arrived at that decision through a rigorous fixed-point calculation, an iterated dominance algorithm, a simple Level-1 heuristic that best-responds to a uniform random prior, or an altruistic effort to maximize joint payoffs. Terminal choice data alone cannot separate these competing hypotheses.
This identification crisis exposed the profound epistemological justification for measuring intermediate information acquisition steps. When human beings evaluate a normal-form game, they do not absorb the entire payoff matrix simultaneously through an instantaneous perceptual intake. Rather, they engage in an active, sequential, and directed information search. The order in which a player inspects payoff cells, the time spent evaluating their own payoffs relative to their opponent’s payoffs, and the structural transitions between strategic coordinates reflect the underlying algorithmic execution of their thought process. If an analyst can observe this intermediate search path, they can open the “black box” of decision-making and measure the actual inputs into the cognitive process.
Consequently, game theory began a fundamental transition from static outcome analysis toward algorithmic and procedural cognitive modeling. Instead of asking exclusively what equilibrium action an agent selects, modern behavioral game theory asks how the agent arrived at that decision. By formalizing human reasoning as an explicit computational sequence—comprising information acquisition, belief formation, payoff comparison, and response execution—researchers gained the tools necessary to differentiate genuine equilibrium reasoning from low-level cognitive heuristics that superficially mimic rational outcomes.
1.3 The Collaborative Agenda of Costa-Gomes, Crawford, and Broseta
This critical intellectual shift crystallized in the late 1990s through the collaboration of Miguel Costa-Gomes, Vincent P. Crawford, and Bruno Broseta. Their research culminated in their landmark 2001 paper published in Econometrica, titled “Cognition and Behavior in Normal-Form Games: An Experimental Study.” At the time, the nascent behavioral literature on bounded rationality in games was heavily fragmented. While several pioneering scholars had begun proposing structural non-equilibrium models, these contributions relied almost entirely on traditional choice data, leaving them vulnerable to charges of ad hoc parameter fitting and non-uniqueness.
The Costa-Gomes, Crawford, and Broseta agenda represented a rigorous fusion of three disparate methodological domains: theoretical econometrics, structural behavioral modeling, and high-precision laboratory experimental protocols. Rather than inventing post-hoc psychological explanations for deviations from Nash equilibrium, the authors set out to construct a unified theoretical taxonomy of decision rules—termed strategic “types”—derived directly from game theory, psychology, and cognitive science. Each decision type was formally articulated not merely as a mapping from games to terminal actions, but as a complete algorithmic procedure with explicit, testable implications for information acquisition.
The primary objective of their 2001 study was to test these structural theories of bounded rationality without relying on ad hoc heuristics or arbitrary curve-fitting techniques. By developing an experimental design that maximized the theoretical separation of decision rules and tracking subjects’ information-gathering processes, they sought to evaluate the cognitive validity of both equilibrium and non-equilibrium models. Their work established the empirical and cognitive foundations for modern behavioral economics, fundamentally redefining how economists operationalize, observe, and evaluate strategic reasoning in laboratory settings.
2. Theoretical Underpinnings: Structural Decision Rules and Level-k Models
2.1 Defining Non-Strategic Baseline Decision Types
To construct a rigorous taxonomy of strategic cognition, Costa-Gomes, Crawford, and Broseta grounded their structural framework in well-defined non-strategic baseline types. In any hierarchical model of mental processing, reasoning must be anchored in an initial, zero-order belief structure. Within this framework, the baseline non-strategic entity is formalised as the Level-0 (L0) type. The L0 type does not engage in strategic deliberation, nor does it attempt to model the payoffs, incentives, or intentions of its counterpart. Instead, an L0 player is operationalized as a uniform random benchmark, selecting among all playable strategies with equal probability. In the authors’ empirical testing, L0 is rarely assumed to exist as a conscious human archetype; rather, it serves as the foundational, internal cognitive anchor that higher-order strategic types employ as their baseline belief regarding opponent behavior.
Alongside this stochastic baseline, the authors formalized non-strategic archetypes governed by social preferences or extreme risk postures rather than strategic calculation. The Altruistic archetype, for instance, evaluates the matrix by maximizing the sum of both players’ payoffs across all possible outcomes, completely disregarding strategic considerations such as individual dominance, incentives to defect, or competitive advantage. The Altruistic type assumes that game play is a collaborative endeavor, seeking Pareto-efficient joint gains regardless of whether the opponent reciprocates.
Conversely, the Pessimistic (or Maximin) archetype operates under an assumption of total vulnerability or extreme risk aversion. A Pessimistic player evaluates each of their own available actions purely by its worst possible outcome, subsequently selecting the action whose minimum payoff is the highest among all alternatives. This archetype requires no understanding of the opponent’s incentives or payoffs; it simply assumes that nature or the opponent will inflict the worst possible state of affairs. The theoretical necessity of anchoring higher-order reasoning in these well-defined non-strategic priors is paramount: without a precisely parameterized, non-strategic baseline, structural models of strategic cognition cannot determine whether an observed choice stems from recursive mentalizing or an unmodeled initial behavioral prior.
2.2 Hierarchical Strategic Types: Level-1 and Level-2 Thinking
Building upon these non-strategic foundations, Costa-Gomes, Crawford, and Broseta operationalized hierarchical strategic decision rules, today broadly recognized as the foundation of Level-k thinking. The primary strategic tier is occupied by the Level-1 (L1) archetype. A Level-1 player is fundamentally strategic in that they actively seek to maximize their expected payoff; however, their cognitive model of the opponent is remarkably basic. An L1 agent assumes that their opponent is a Level-0 non-strategic entity whose actions are uniformly distributed across the strategy space. Consequently, an L1 player calculates the arithmetic mean of their own payoffs across each column for every row available to them, choosing the action that yields the highest expected value against this uniform random distribution. The L1 player exhibits strategic intent, yet their mental model of the other player contains zero strategic content.
The hierarchy ascends with the Level-2 (L2) archetype, which embodies a richer, counterfactual model of the other player. An L2 player recognizes that the other agent is an optimizing strategic actor, but attributes to that agent an L1 cognitive architecture. Thus, an L2 player assumes their opponent will play the action that is optimal against a uniform random prior. The L2 decision process unfolds in two distinct mental steps: first, the player inspects the opponent’s payoffs across the matrix, calculates the column that maximizes the opponent’s expected payoff against a uniform prior, and determines the opponent’s anticipated L1 choice. Second, the L2 player inspects their own payoffs within that specific predicted column and selects their best response.
The iterative depth of reasoning embedded within these types provides an empirical parameterization of individual cognitive processing limits. Rather than assuming infinite cognitive recursion, the Level-k framework posits that human working memory and mentalizing capacity are strictly constrained, with individuals rarely ascending beyond one or two steps of recursive induction. A critical theoretical distinction exists between these structural Level-k archetypes and classical epistemic hierarchies. Whereas classical epistemic game theory constructs infinite hierarchies of mutual beliefs about rationality, Level-k models introduce a structural cognitive anchoring effect: each level constructs a myopic, non-reciprocal belief that the opponent operates precisely one level below themselves, eliminating the computational necessity of fixed-point calculations.
2.3 Equilibrium Types: Iterated Dominance and Nash Archetypes
In direct contrast to bounded Level-k heuristics, Costa-Gomes, Crawford, and Broseta formalized classical game-theoretic constructs as distinct behavioral types. The simplest of these are the iterated dominance archetypes, which evaluate games through sequential elimination of inferior strategies rather than immediate fixed-point calculations. The Dominance-1 (D1) type executes precisely one round of strict dominance deletion. A D1 agent scans their opponent’s available strategies to determine whether any action yields strictly lower payoffs than another action across all possible responses. If a strictly dominated strategy is detected, the D1 player deletes it from consideration, assumes the opponent will never select it, and subsequently maximizes expected payoff over the reduced game using a uniform random belief over the remaining options.
The Dominance-2 (D2) archetype pushes this logic one step further by executing two sequential rounds of iterated strict dominance calculations. A D2 agent systematically deletes the opponent’s strictly dominated strategies, deletes their own strictly dominated strategies in the reduced game, and then identifies their optimal response. The D1 and D2 archetypes represent structural, algorithmic approximations of hyper-rationality that stop short of full fixed-point analysis, offering an empirical bridge between naive heuristics and classical game theory.
The pinnacle of standard normative theory is embodied by the Nash Equilibrium archetype. A pure Nash type evaluates the game by identifying strategy profiles that constitute mutual best responses across all states. In games with a unique pure-strategy Nash equilibrium, the Nash type identifies and executes the strategy mandated by this equilibrium coordinate. The conceptual differences between these formulations are profound:
- Algorithmic Complexity: Iterative elimination (D1, D2) relies on local, pairwise comparisons of payoff vectors to discard dominated strategies.
- Fixed-Point Computation: The Nash archetype requires solving a simultaneous system of mutual optimization conditions, demanding that beliefs and actions be verified concurrently across all matrix cells.
- Cognitive Tractability: Human working memory handles sequential elimination far more naturally than the non-linear, bidirectional cognitive loops necessary to locate fixed points in complex normal-form matrices.
2.4 The Sophisticated Decision Type
Recognizing the potential gap between theoretical equilibrium perfection and empirical meta-cognition, Costa-Gomes, Crawford, and Broseta introduced the Sophisticated decision archetype. In classical game theory, a rational player assumes all other players are equally rational and conform to equilibrium. In contrast, the Sophisticated player is conceptualized as an optimal forecaster who possesses rational expectations regarding the actual empirical distribution of decision types within the player population. A Sophisticated agent does not assume that opponents are textbook Nash players, nor do they assume opponents are exclusively L1 or L2 actors; rather, they form an accurate probability distribution over the entire population mixture—comprising naive, boundedly rational, and equilibrium types.
Formally, let the empirical population be characterized by a probability vector $P = (p_{\text{L1}}, p_{\text{L2}}, p_{\text{D1}}, p_{\text{Nash}}, dots)$, where each $p_j$ represents the true proportion of type $j$ agents in the experimental pool. A Sophisticated player selects an action $a^*$ that maximizes expected utility against this empirical mixture:
$$a^* = arg\max_{a_i in A_i} \sum_{j} p_j \cdot u_i(a_i, a_{-i}^j)$$
where $a_{-i}^j$ denotes the action predicted for a player of type $j$, and $u_i$ is the player’s payoff function. This archetype establishes a vital epistemological distinction between theoretical equilibrium perfection and empirical meta-cognition. A player may possess deep cognitive sophistication, but if the surrounding population is dominated by naive Level-1 agents, the optimal strategic response will systematically deviate from the Nash equilibrium to exploit those naive tendencies.
However, the information processing demands required to operationalize the Sophisticated archetype are immense. To execute this rule, a subject must not only possess the cognitive capacity to compute the actions of all lower-order types (L1, L2, D1, Nash), but must also form accurate Bayesian priors regarding the population proportions of each type. Consequently, while the Sophisticated type represents the theoretical ideal of an economically rational agent operating under behavioral realism, its sheer computational burden raises profound questions regarding its empirical prevalence among human subjects.
3. Experimental Architecture: Design of the 18 Normal-Form Games
3.1 Matrix Construction and Strategic Dominance Configurations
The experimental foundation of the Costa-Gomes, Crawford, and Broseta study rests upon an eighteen-game panel of two-player, normal-form strategic matrices. Rather than presenting subjects with standardized, symmetric textbook games such as the Prisoner’s Dilemma or Battle of the Sexes, the authors engineered a structurally diverse sequence of matrices designed to probe distinct dimensions of strategic reasoning. The games varied systematically in dimension—including 2×2, 3×2, 2×3, 4×2, and 3×3 configurations—preventing subjects from relying on visual symmetry or memorized heuristic layouts.
Crucially, the games were divided along rigorous theoretical lines regarding their strategic solvability:
- Dominance-Solvable Games: Games solvable by iterated deletion of strictly dominated strategies, requiring between one to three rounds of elimination to yield a unique solution.
- Equilibrium-Dependent Games: Games that contained no strictly dominated strategies for at least one player, thereby requiring simultaneous, fixed-point Nash reasoning to resolve.
The authors also varied the presence of weakly and strictly dominant strategies between Row and Column roles, ensuring that subjects encountered matrices where their own role possessed a dominant strategy while the opponent did not, and vice versa. Payoffs were calibrated to provide meaningful financial incentives while mitigating the confounding effects of risk aversion. Cell payoffs represented lottery tickets under a binary lottery procedure, converting expected game payoffs directly into linear probabilities of winning cash prizes, thereby theoretically neutralizing risk-averse inclinations under expected utility theory.
3.2 Strategic Separation of Theoretical Decision Rules
The central methodological innovation of the authors’ experimental architecture was the strategic separation of theoretical decision rules. In standard matrix games, different decision types often prescribe identical actions. For instance, if an action is both strictly dominant and part of a unique Nash equilibrium, observing a subject choose that action yields zero diagnostic power regarding whether the choice was driven by a simple Level-1 heuristic, iterated dominance, or Nash calculation. To resolve this identification problem, Costa-Gomes, Crawford, and Broseta designed their payoff matrices to maximize behavioral divergence among rival types.
Across the eighteen games, the payoff values were calibrated to ensure that candidate decision types—specifically Altruistic, Pessimistic, L1, L2, D1, D2, Nash, and Sophisticated—predicted distinct, orthogonal action profiles across the sequence of matrices. The authors evaluated this design feature using a formal separation index, systematically minimizing the correlation between the action vectors predicted by different decision rules.
While complete separation in every single game is mathematically impossible in low-dimensional matrices, the sequence of eighteen games was engineered such that no two competing decision types shared identical predictions across more than a small fraction of the games. By confronting subjects with this panel, the experimental design guaranteed that an agent’s aggregate vector of eighteen terminal choices could serve as a unique behavioral fingerprint, decisively isolating their underlying cognitive model from competing alternatives.
3.3 Control Mechanisms: Asymmetry and Opponent Anonymity
To ensure that information acquisition and choice data reflected strategic reasoning rather than behavioral confounds, the experimental architecture incorporated strict structural controls. First, the authors enforced strict payoff asymmetry across all eighteen games. In symmetric games, players can rely on focal points, fairness heuristics, or structural visual cues (such as identifying the main diagonal) without engaging in strategic mentalizing. Asymmetric payoffs eliminated these confounding focal points, forcing subjects to evaluate the opponent’s unique incentive structure if they wished to anticipate their opponent’s play.
Second, the experimental protocol eliminated repeated-game dynamics, reputation formation, and reciprocity through an anonymous, single-shot random matching design. Subjects played each game exactly once without receiving intermediate feedback. After each matrix choice, the system advanced to the next game without revealing the partner’s selection or the resulting payoff. This protocol severed the multi-period dynamics that typically generate social preferences, retaliatory punishments, or reinforcement learning, isolating purely introspective, initial-response strategic reasoning.
Finally, the games were presented using neutral framing protocols. Instructions were articulated entirely in abstract terms (e.g., “Row Player,” “Column Player,” “Option 1,” “Box A”), avoiding loaded terminology like “opponent,” “partner,” “win,” or “lose.” By scrubbing conversational and situational framing biases from the experimental environment, the authors ensured that observed choices were driven solely by the strategic mechanics of the payoffs rather than psychological framing effects.
4. Innovative Methodology: Information Search Tracking via MouseLab
4.1 Architecture of the MouseLab Software Interface
To directly observe the procedural, algorithmic operations of subjects as they contemplated their decisions, Costa-Gomes, Crawford, and Broseta deployed the MouseLab software interface. In an unconstrained experimental setting, a subject views a complete normal-form matrix displayed openly on a screen, leaving the researcher unable to determine which payoff cells are being processed at any given moment. MouseLab circumvented this limitation by concealing every payoff value inside an opaque visual box labeled with the identity of the player and coordinate.
To inspect a payoff, the subject was required to move their computer cursor over the designated box. The box remained open and visible only while the mouse cursor remained inside its boundaries; the moment the cursor moved outside, the box immediately closed, returning the payoff value to an occluded state. The interface enforced a strict single-cell opening constraint: only one box could be opened at any instant in time. Through this mechanism, the MouseLab system continuously logged high-resolution process data:
- The exact chronological sequence of payoff cell openings.
- The precise duration of exposure for each individual inspection (measured in milliseconds).
- The frequency and trajectory of revisitation patterns across strategic coordinates.
This experimental design eliminated visual periphery confounds, ensuring that subjects could not acquire information via peripheral vision. At the same time, the authors carefully calibrated the mouse-tracking interface to minimize cognitive friction, ensuring that manual cursor movement did not impede natural cognitive processing or distort the underlying decision-making algorithms.
4.2 Formulating Procedural Search Hypotheses
The primary methodological breakthrough of tracking information search via MouseLab was the formulation of ex ante, procedural search hypotheses. Rather than treating gaze paths as unstructured descriptive data, Costa-Gomes, Crawford, and Broseta derived necessary and sufficient search sequences directly from the mathematical definitions of the strategic decision types. If a decision maker executes a specific algorithm, that algorithm logically requires access to a precise subset of the payoff matrix before a rational choice can be made.
These algorithmic derivations generated distinct directional search profiles. To evaluate an action under a Level-1 heuristic, an agent needs access only to their own payoffs across different states; the opponent’s payoffs are mathematically irrelevant. Consequently, an L1 thinker should exhibit a predominance of vertical transitions (comparing their own payoffs within an action across columns). Conversely, to determine an opponent’s Level-1 response, a Level-2 player must first inspect the opponent’s payoffs horizontally, locate the opponent’s maximum expected return, and only then shift attention to their own payoffs within the predicted column.
To formalize these requirements, the authors defined explicit search compliance criteria:
- Minimal Relevant Information Sets: The specific subset of payoff cells that a given decision type must inspect to compute their optimal action.
- Directional Search Transitions: The structural movement of attention between cells—such as intra-action comparisons versus inter-player cross-checks—required by a given algorithm.
- Procedural Compliance: A measure evaluating whether a subject’s observed search path meets the necessary informational prerequisites of a candidate decision type, evaluated independently of their terminal choice.
These criteria decoupled cognitive process from behavioral outcome, allowing the researchers to test whether a player who made a seemingly sophisticated choice actually acquired the information required to calculate that choice strategically.
4.3 Quantitative Metrics for Search Analysis
Transforming continuous MouseLab logs into rigorous econometric data required developing a suite of standardized quantitative metrics. The raw data—consisting of time-stamped cursor coordinates—was distilled into structural indicators capturing the distribution, sequence, and density of cognitive attention. The first fundamental metric was absolute inspection duration, which aggregated the total time a subject spent examining specific payoff subsets (own payoffs vs. opponent payoffs), providing a direct proxy for cognitive load and attention allocation.
To analyze sequential dynamics, the authors constructed transition matrices. By categorizing cell transitions into horizontal (scanning across columns), vertical (scanning down rows), and diagonal movements, they quantified the relative frequency of intra-box versus inter-box gaze shifts. Crucially, they differentiated between transitions that compared payoffs belonging to the same player and those that cross-checked outcomes between both players. These transitions were formalized into structural metrics such as the Occurrence metric (which evaluates whether every cell in a type’s minimal information set was inspected at least once) and the Adjacency metric (which measures whether relevant cell inspections occurred consecutively in the required algorithmic order).
Finally, the authors confronted substantial individual-level baseline heterogeneity in motor execution, mouse dexterity, and baseline reading speeds. A raw lookup duration of 500 milliseconds might reflect deep cognitive deliberation for one subject, but an unthinking pause for another. To address this variance, the researchers normalized their search metrics relative to each subject’s aggregate processing profile across the entire panel of games. By evaluating relative gaze proportions, normalized transition densities, and structural compliance ratios, the econometric framework isolated genuine cognitive deliberation from idiosyncratic motor variance.
5. Mapping Cognitive Search Signatures to Behavioral Types
5.1 Information Sets for Naive and Heuristic Archetypes
The structural theories of bounded rationality formulated by Costa-Gomes, Crawford, and Broseta generate sharply divergent cognitive search signatures across candidate decision types. For the naive and heuristic archetypes, these search requirements are strikingly localized and minimalistic, reflecting their computationally myopic decision rules.
The Level-1 (L1) heuristic, which assumes an opponent randomizes uniformly over playable actions, requires information strictly concerning the player’s own payoffs. To compute the expected utility of any available action, an L1 player needs only to sum their own payoffs across each column and divide by the number of columns. Opponent payoffs are entirely extraneous to this calculation. Consequently, the necessary information search signature for a pure Level-1 thinker is characterized by an exclusive concentration on self-payoff cells and a near-total absence of gaze fixations on opponent payoff boxes. A subject whose search trace displays frequent, prolonged inspections of the opponent’s payoffs violates the fundamental procedural mechanics of the L1 algorithm.
Similarly, the Pessimistic (Maximin) archetype requires a hyper-localized search footprint. To identify the maximin choice, a player must scan each row of their own payoffs, record the minimum entry in each row, and select the row containing the highest minimum. This rule requires zero attention to opponent payoffs and only selective attention to self-payoffs. In contrast, the Altruistic archetype generates an entirely different search profile: it demands an exhaustive, balanced inspection of both self-payoffs and opponent-payoffs across every cell, followed by explicit cross-cell additions to evaluate joint outcomes. These distinct informational footprints allow researchers to differentiate heuristic archetypes that might otherwise generate identical choices in specific strategic matrices.
5.2 Search Complexities of Higher-Level Strategic Types
In stark contrast to naive heuristics, higher-level strategic types generate complex, structured information search signatures that reflect counterfactual mentalizing and multi-step induction. The Level-2 (L2) decision algorithm requires a sophisticated, non-linear search sequence that mirrors its two-step cognitive architecture. Before an L2 player can evaluate their own optimal move, they must first identify the optimal move of an L1 opponent. Consequently, the initial phase of an L2 search is characterized by an extensive scan of the opponent’s payoff matrix, calculating the column that maximizes the opponent’s average returns.
Once this counterfactual computation is complete, the L2 search signature shifts abruptly: the player transitions their gaze to their own payoff cells, focusing selectively on the specific column predicted to be chosen by the L1 opponent. This creates a distinct, observable cognitive profile:
- Asymmetric Payoff Prioritization: Early lookups are dominated by opponent-payoff boxes, often displaying higher initial dwell times than own-payoff boxes.
- Interleaved Cell Transitions: Gaze sequences shift systematically from opponent-payoff aggregations to own-payoff evaluations.
- Elevated Cognitive Load: Total dwell times and revisitation rates are significantly higher than those observed among L1 or Maximin players, reflecting working-memory storage of intermediate calculations.
This sequential asymmetry provides an unmistakable procedural footprint. A player cannot be genuinely operating as a Level-2 thinker if they fail to systematically inspect the opponent’s payoff matrix prior to settling on their own chosen action.
5.3 Cognitive Profiles of Dominance and Equilibrium Solvers
The search requirements escalate dramatically when moving from bounded Level-k heuristics to iterated dominance and fixed-point equilibrium solvers. For the Dominance-1 (D1) and Dominance-2 (D2) archetypes, the search signature must reflect the systematic, pairwise elimination of dominated options. To detect whether an opponent’s strategy is strictly dominated, a player must compare the opponent’s payoffs across every single row for two parallel columns. This mandates a dense pattern of horizontal, intra-opponent payoff comparisons. Only after strict dominance is confirmed can the search focus shift to selecting an optimal response within the reduced matrix.
The search profile demanded by the pure Nash Equilibrium archetype is the most complex of all. To verify a Nash equilibrium in a general normal-form game lacking dominant strategies, an agent cannot rely on one-directional sweeps. Instead, the player must execute exhaustive, bidirectional cell transitions:
- Inspect cell $(i, j)$ and verify that own payoff $u_1(i, j)$ is maximal given column $j$.
- Simultaneously inspect cell $(i, j)$ to confirm that opponent payoff $u_2(i, j)$ is maximal given row $i$.
- Repeat this cross-checking verification across candidate cells until a stable coordinate pair is located.
This fixed-point search imposes immense demands on human working memory. Subjects must simultaneously hold multiple payoff coordinates in active memory while verifying mutual optimality conditions across bidirectional axes. When human subjects attempt this process under experimental conditions, MouseLab traces frequently capture chaotic, fragmented search paths marked by repeated revisitations and prolonged dwell times, reflecting the structural cognitive bottlenecks inherent in computing fixed-point solutions in normal-form games.
6. Econometric Framework for Joint Choice and Search Analysis
6.1 The Choice-Based Maximum Likelihood Specification
To translate behavioral observations and process-tracing metrics into statistical inferences, Costa-Gomes, Crawford, and Broseta formulated a rigorous structural econometric framework. The foundation of this empirical methodology is a choice-based maximum likelihood specification. The authors assumed that the experimental subject population is comprised of a discrete set of latent strategic types, denoted $k in \mathcal{K}$, where $\mathcal{K} = {\text{L1}, \text{L2}, \text{D1}, \text{D2}, \text{Nash}, \text{Sophisticated}, dots}$. Each latent type $k$ prescribes an optimal, deterministic action $c_{ik}^*$ for subject $i$ in game $g in {1, dots, 18}$.
To accommodate behavioral noise, the econometric model incorporates an explicit strategic error parameter. Following the logic of the trembling-hand formulation, a subject of type $k$ is assumed to choose the type-prescribed action $c_{ik}^*$ with probability $(1 – \epsilon_k)$, where $\epsilon_k in [0, 1]$ represents the error rate. In the event of an error (a “tremble”), the subject randomizes uniformly over all $J_g$ available strategies in game $g$. Thus, the probability of observing action $c_{ig}$ conditional on the subject being of latent type $k$ is defined as:
$$P(c_{ig} mid k, \epsilon_k) = (1 – \epsilon_k) \cdot \mathbb{I}(c_{ig} = c_{ik}^*) + \frac{\epsilon_k}{J_g}$$
Assuming independence of errors across the panel of eighteen games conditional on type, the overall likelihood of observing a subject’s complete eighteen-element action vector $C_i = (c_{i1}, dots, c_{i18})$ conditional on type $k$ is the product of individual game probabilities:
$$L_i(C_i mid k, \epsilon_k) = \prod_{g=1}^{18} P(c_{ig} mid k, \epsilon_k)$$
At the aggregate population level, the unconditional likelihood of the sample is obtained by modeling the population as a mixture of types, where $p_k$ represents the latent proportion of type $k$ within the population, satisfying $\sum_{k in \mathcal{K}} p_k = 1$. The aggregate log-likelihood is then maximized across subjects and types, utilizing standard information criteria—such as the Akaike Information Criterion (AIC) and the Bayesian Information Criterion (BIC)—to evaluate model fit, penalize over-parameterization, and establish the statistical dominance of structural type models over naive random benchmarks.
6.2 Integrating Information Search into the Likelihood Formulation
While choice-based maximum likelihood provides a robust structural framework, it remains vulnerable to the fundamental identification problem when distinct types predict identical choices. The transformative contribution of Costa-Gomes, Crawford, and Broseta was the direct integration of intermediate information search data into the structural likelihood function. Rather than treating MouseLab search logs as an informal diagnostic check, they formalized search metrics as endogenous, probabilistic indicators of latent type identity.
Let $M_{ig}$ denote the observed information acquisition vector (comprising dwell times, cell lookups, and transition sequences) generated by subject $i$ in game $g$. For each theoretical type $k$, the authors defined explicit, binary or continuous search compliance criteria, denoted $S_{ikg} in [0, 1]$, which measured the degree to which $M_{ig}$ satisfied the necessary and sufficient information set required by type $k$. To model compliance probabilistically, the authors parameterized procedural errors: a subject of type $k$ satisfies type $k$‘s required search pattern with probability $(1 – \zeta_k)$, and exhibits an arbitrary, non-compliant search pattern with error probability $\zeta_k$.
The joint probability of observing both the terminal choice $c_{ig}$ and the information search pattern $M_{ig}$, conditional on the subject belonging to latent type $k$, is formulated as the product of the conditional choice and search distributions:
$$P(c_{ig}, M_{ig} mid k, \epsilon_k, \zeta_k) = P(c_{ig} mid k, \epsilon_k) \cdot P(M_{ig} mid k, \zeta_k)$$
By compounding these joint probabilities across all eighteen games, the joint likelihood function for subject $i$ becomes:
$$L_i(C_i, M_i mid k, \epsilon_k, \zeta_k) = \prod_{g=1}^{18} P(c_{ig} mid k, \epsilon_k) \cdot P(M_{ig} mid k, \zeta_k)$$
This formulation resolved severe identification degeneracies. If subject $i$‘s terminal choice vector was equally consistent with both Level-2 and Nash predictions, the joint likelihood evaluated the compliance of $M_i$ with the respective search requirements of L2 and Nash. If the subject systematically failed to inspect the cell transitions mandatory for Nash verification while consistently exhibiting the opponent-payoff scanning signature of Level-2, the joint econometric estimation unequivocally assigned the subject to the Level-2 archetype.
6.3 Identification and Type-Classification Robustness
To validate the stability and diagnostic reliability of their joint econometric framework, Costa-Gomes, Crawford, and Broseta conducted rigorous identification and sensitivity analyses. A primary concern was whether individual type classifications were stable across the panel of eighteen distinct games, or whether subjects experienced rapid learning, type-switching, or fatigue as the experiment progressed. By splitting the panel into chronological halves and estimating separate type parameters across game blocks, the authors demonstrated remarkable within-subject type stability: an individual classified as a Level-1 thinker in early games rarely transformed into a Level-2 or Nash solver in later rounds, confirming that types reflect stable, cognitive processing traits rather than volatile experimental noise.
To assess the false-positive diagnostic rate of their joint choice-search likelihood estimator, the authors performed extensive Monte Carlo simulations using synthetic, randomized pseudo-agents. Synthetic agents were generated to play games by selecting actions and search sequences uniformly at random or through simulated trembling-hand algorithms. The econometric estimation procedure was then applied to these synthetic datasets. The simulations demonstrated that the combined choice-and-search likelihood had an exceptionally low false-positive rate: pure random noise was virtually never misclassified as an L2, D1, or Nash type, because the joint probability of accidentally selecting both the correct action vector and the correct information search transitions across eighteen asymmetric matrices is vanishingly small.
The framework also established clear boundaries for unclassifiable or hybrid-reasoning subjects. Individuals whose behavioral vectors exhibited elevated error parameters ($\epsilon_i > 0.5$) or whose search compliance fell below statistical thresholds were assigned to an unclassified or random residual category. Rather than forcing every subject into a theoretical box, the joint econometric model distinguished clearly between subjects displaying systematic bounded rationality and those whose cognitive deliberation was dominated by procedural noise.
7. Empirical Findings: The Actual Distribution of Strategic Types
7.1 The Preponderance of Bounded Strategic Rationality
The empirical estimation of the Costa-Gomes, Crawford, and Broseta dataset yielded a profound revelation regarding the nature of human strategic cognition: actual human populations are heavily concentrated in low-level, boundedly rational cognitive archetypes. Across multiple econometric specifications, the overwhelming majority of subjects were classified as either Level-1 (Naive) or Level-2 thinkers. Together, these two boundedly rational types accounted for between 65% and 80% of the classifiable subject pool, demonstrating that human strategic reasoning rarely extends beyond two steps of mental iteration.
| Strategic Type | Theoretical Core Assumptions | Search Footprint Signature | Observed Sample Proportion |
|---|---|---|---|
| Level-1 (L1) | Best-responds to uniform random non-strategic prior (L0) | Exclusively self-payoff cells; ignores opponent payoffs | ~45% – 50% |
| Level-2 (L2) | Best-responds to an anticipated Level-1 opponent | Scans opponent payoffs first, then checks own payoffs | ~25% – 30% |
| Dominance-1 (D1) | Deletes strictly dominated strategies (one iteration) | Systematic pairwise comparisons of opponent payoffs | ~5% – 8% |
| Nash Equilibrium | Fixed-point mutual best responses across all states | Exhaustive bidirectional lookups and cross-checks | < 5% |
| Sophisticated | Best-responds to empirical population mixture | Extensive, multi-type cognitive search coverage | 0% (Virtually undetected) |
This distribution highlights an immense divergence between theoretical game-theoretic equilibrium predictions and actual human deliberation. While subjects in experimental settings are actively attempting to maximize their monetary rewards, their cognitive machinery does not evaluate strategic games via global equilibrium balance. Level-1 players act with clear economic purpose, but they fail to mentalize their opponent’s incentives entirely. Level-2 players take their opponent’s payoffs into account, but model their counterpart as a simple, non-strategic actor. The finding that human strategic deliberation operates almost exclusively within this shallow cognitive hierarchy has been replicated across dozens of subsequent experimental studies, establishing the Level-k distribution as an empirical regularity in unlearned, normal-form games.
7.2 The Rarity of Pure Equilibrium and Dominance Solvers
Perhaps the most striking and consequential finding of the 2001 study was the near-total absence of subjects who behaved as classical equilibrium solvers. Despite being recruited from elite student populations at major research universities, less than five percent of subjects exhibited behavior and search patterns consistent with the pure Nash Equilibrium type. When confronting complex, asymmetric matrices lacking dominant strategies, human subjects almost never deploy the simultaneous, fixed-point reasoning required by classical game theory.
Even more surprising was the empirical rarity of iterated dominance types (D1 and D2). One of the most pervasive assumptions in neoclassical economics is that rational actors will recognize and eliminate strictly dominated strategies. However, in the Costa-Gomes et al. experiment, subjects executing even a single round of iterated strict dominance represented a minor fraction of the population, while D2 types—requiring two iterations of dominance deletion—were virtually undetectable. The cognitive penalty, working-memory load, and procedural failure rates associated with executing multi-step matrix elimination under experimental conditions proved to be a prohibitive barrier for most subjects.
The broader implications of this finding for theoretical economics are severe. A vast array of modern microeconomic models—including advanced auction designs, dynamic bargaining protocols, voting mechanisms, and market signaling games—rely on the assumption of iterated dominance or Nash equilibrium solvability under common knowledge of rationality. The empirical evidence demonstrating that human agents rarely execute even two rounds of dominance deletion suggests that many textbook economic mechanisms rest on descriptive foundations that do not hold in initial strategic interactions.
7.3 The Puzzle of the Sophisticated Type
In theoretical formulations of bounded rationality, the Sophisticated type represents the pinnacle of strategic intelligence: an agent who accurately estimates the empirical mixture of naive, intermediate, and rational agents across the population and plays the optimal empirical best response. In theory, such an agent should systematically out-earn all other participants in the experimental economy. Yet, in the empirical estimation conducted by Costa-Gomes, Crawford, and Broseta, the estimated proportion of Sophisticated players was essentially zero.
This puzzle stems from two distinct factors. The first is computational intractability: to operate as a genuine Sophisticated type, an individual must compute the optimal actions for multiple structural types (L1, L2, D1, Nash) across every game, assign accurate probability weights to each type based on aggregate population behavior, and sum the expected payoffs across this complex distribution. This represents a monumental cognitive task that far outstrips the working-memory capacity of human decision-makers.
The second factor is strategic convergence. In many normal-form games, the action that best-responds to an empirical population mixture heavily dominated by Level-1 agents is mathematically identical to the action prescribed by the Level-2 model. Because an L2 player best-responds to an L1 agent, and the real population is dominated by L1 agents, an agent who merely acts as an L2 thinker will achieve near-optimal payoffs without needing to perform the elaborate meta-cognitive calculations of the Sophisticated type. Consequently, from both a cognitive load perspective and an evolutionary payoff standpoint, the Level-2 heuristic acts as an efficient, highly successful behavioral proxy for true sophistication, rendering the emergence of pure Sophisticated types unnecessary in practice.
8. Choice Data versus Process Data: Validation and Divergence
8.1 Concordance Between Information Acquisition and Chosen Moves
The ultimate methodological validation of the Costa-Gomes, Crawford, and Broseta paradigm lies in the remarkable concordance observed between subjects’ information acquisition patterns and their final terminal choices. In an experimental environment, it is theoretically possible that process-tracing data might capture random eye movements, idle cursor wandering, or superficial visual distractions unrelated to the subject’s actual decision rule. However, the econometric findings revealed an exceptionally high degree of internal consistency between intermediate search paths and ultimate choices.
For the majority of subjects classified as Level-1 thinkers based on their action vectors, their MouseLab logs revealed an almost complete procedural neglect of opponent payoff boxes. These subjects spent the vast majority of their lookup durations examining self-payoff coordinates, systematically ignoring the numbers that determined their partner’s welfare. Conversely, subjects classified as Level-2 thinkers exhibited high procedural compliance with the two-stage search sequence mandated by the L2 model, consistently inspecting the opponent’s payoffs before focusing on their own choice. This alignment between independent behavioral channels—cursor movement sequences on one hand and discrete button clicks on the other—provided strong econometric confirmation that structural Level-k models capture real, operational cognitive procedures rather than mathematical artifacts.
8.2 Unmasking Spurious Classifications Through Information Search
While concordance affirmed the validity of bounded rationality models, the true diagnostic value of information search data emerged in instances of deep behavioral divergence. In several critical matrix configurations, the action predicted by the textbook Nash equilibrium coincided with the action prescribed by a low-level heuristic, such as a Level-1 or Maximin rule. When evaluated strictly under choice-based maximum likelihood, subjects playing these games were frequently categorized as Nash or Dominance-1 types, leading to an overly optimistic assessment of population rationality.
However, when the authors integrated MouseLab process data into the joint likelihood framework, these spurious equilibrium classifications collapsed. The search logs revealed that many subjects who selected the “rational” Nash action had never opened the opponent payoff cells necessary to compute a Nash equilibrium or verify iterated dominance. In some cases, subjects selected the unique equilibrium move without inspecting their opponent’s payoffs for even a fraction of a second. By unmasking these instances of accidental conformity, the process data revealed that what appeared to be sophisticated equilibrium play was actually the byproduct of naive heuristics that happened to coincide with equilibrium predictions. Without intermediate search tracking, classical economics would have erroneously concluded that these subjects were executing hyper-rational fixed-point calculations.
8.3 Procedural Noise and Heuristic Approximations
Alongside its diagnostic successes, the tracking of cognitive search paths revealed the presence of pervasive procedural noise and heuristic approximations that diverge from idealized algorithmic models. Real human reasoning does not unfold with the mechanical perfection of a computer algorithm. Even subjects firmly classified within a specific strategic type exhibited deviations, including accidental lookups, brief lapses in attention, and idiosyncratic scanning behaviors.
A major source of procedural divergence was working-memory retention. In multiple-round experiments, subjects frequently retained payoff values in memory across sequential inspections, eliminating the need to repeatedly reopen previously viewed boxes. An observer tracking mouse clicks might see an incomplete search sequence and conclude the player lacked sufficient information, when in reality the subject had already stored the relevant numbers in short-term memory during an earlier inspection. Furthermore, subjects exhibited substantial individual heterogeneity in spatial scanning routines, with some players preferring systematic top-to-bottom sweeps before engaging in directed calculation. Resolving the tension between the pristine, rigid definitions of theoretical algorithms and the noisy, organic realities of human biological cognition remains one of the core empirical challenges highlighted by the Costa-Gomes et al. methodology.
9. Methodological Implications for Behavioral and Experimental Economics
9.1 Challenging the Revealed-Preference Orthodoxy
The epistemological fallout of the Costa-Gomes, Crawford, and Broseta investigation extended far beyond the boundaries of normal-form game theory; it struck at the heart of the twentieth-century neoclassical methodology. Since the publication of Milton Friedman’s 1953 essay, “The Methodology of Positive Economics,” mainstream economics had largely embraced an instrumentalist “as-if” doctrine. Friedman argued that the realism of an economic model’s behavioral assumptions—whether agents actually calculate marginal costs, compute fixed points, or execute dynamic programming—is completely irrelevant. So long as the model yields accurate predictions of market outcomes, economists were licensed to treat agents as if they were hyper-rational optimizers, viewing the human mind as an impenetrable black box.
The 2001 study delivered a decisive blow to this methodological stance. By proving that internal cognitive operations can be systematically operationalized, tracked, and structurally estimated using high-resolution process data, the authors demonstrated that the “as-if” assumption is an unnecessary methodological handicap. The study showed that treating boundedly rational agents “as if” they were Nash players leads to systemic misattributions, erroneous policy forecasts, and a total failure to predict behavior in novel, unlearned environments. By establishing intermediate process measures—such as lookup sequences, dwell times, and attention distributions—as valid, rigorous economic data, Costa-Gomes, Crawford, and Broseta helped shift behavioral economics from post-hoc curve fitting toward genuine algorithmic realism.
9.2 Methodological Rigor in Laboratory Search Protocols
The introduction of MouseLab into game-theoretic experiments also prompted intense methodological debate regarding the design of laboratory search protocols. A primary concern centered on potential demand characteristics and artificial friction introduced by the user interface. Opening payoff boxes with a mouse cursor imposes a mechanical transaction cost that does not exist when reviewing a traditional printed matrix. Skeptics questioned whether this physical friction might distort subjects’ natural strategic deliberation, potentially inducing more simplistic, heuristic behavior than would occur under unconstrained visual inspection.
To address these concerns, the authors established rigorous experimental controls that have become standard across experimental economics:
- Pre-Experimental Interface Training: Subjects completed extensive neutral training exercises to master mouse mechanics and minimize interface-induced latency before playing any strategic games.
- Minimizing Motor Costs: Payoff boxes were designed with large target areas and zero opening lag to reduce physical strain and manual errors.
- Occlusion Consistency: The single-box opening protocol was held completely constant across all games, ensuring that search frictions could not explain differential strategic behavior across matrix types.
Subsequent external validity studies confirmed that while manual search interfaces introduce a minor latency baseline, the structural distribution of strategic types and the relative allocation of attention between own and opponent payoffs remain remarkably invariant across diverse tracking protocols.
9.3 The Structural Econometric Approach to Bounded Rationality
Beyond its substantive findings, the Costa-Gomes, Crawford, and Broseta paper provided a gold standard for structural econometrics applied to behavioral data. Before its publication, experimental studies of bounded rationality were frequently criticized for lacking mathematical rigor, often relying on informal descriptive statistics, post-hoc verbal categorization, or unrestricted parameter estimation that risked overfitting the data.
The authors countered this criticism by establishing a disciplined structural methodology. Rather than letting parameters vary freely, they specified candidate behavioral types and their associated search requirements strictly ex ante, derived directly from mathematical theory. By incorporating explicit error distributions ($\epsilon_k$ and $\zeta_k$) into a unified likelihood framework, they allowed the data to probabilistically select the best-fitting models while penalizing unnecessary complexity via formal information criteria. This structural econometric approach proved that behavioral economics could study bounded rationality and psychological limits with the same analytical rigor, mathematical clarity, and econometric discipline that characterized neoclassical economics.
10. Comparative Analysis: Alternative Frameworks and Successor Models
10.1 Comparison with Stahl and Wilson’s Discrete Type Models
The intellectual trajectory that culminated in the 2001 study was significantly shaped by the pioneering work of Dale Stahl and Paul Wilson, who published foundational discrete-type econometric analyses of strategic thinking in 1994 and 1995. Stahl and Wilson were among the first economists to formalize a hierarchical model of bounded rationality, estimating the proportions of Level-0, Level-1, and Level-2 thinkers using experimental choice data from symmetric 3×3 normal-form games.
However, the Costa-Gomes, Crawford, and Broseta framework diverged from and improved upon the Stahl and Wilson approach in three critical dimensions:
- Experimental Diagnostic Separation: Stahl and Wilson utilized symmetric normal-form games where theoretical decision rules often generated overlapping action predictions, creating severe identification ambiguities. Costa-Gomes et al. engineered an eighteen-game asymmetric panel that maximized the orthogonal separation of candidate types.
- Choice-Only vs. Joint Choice-Search Estimation: Stahl and Wilson relied entirely on revealed terminal choices, leaving their estimations vulnerable to accidental heuristic mimicry. Costa-Gomes et al. integrated MouseLab process data, introducing an independent empirical channel to verify cognitive mechanisms.
- Treatment of Unobserved Heterogeneity: While Stahl and Wilson estimated complex models that included continuous mixtures and idiosyncratic belief parameters, Costa-Gomes et al. imposed strict ex ante discipline on their decision rules, preventing the over-fitting of strategic types.
Consequently, while Stahl and Wilson demonstrated the existence of strategic hierarchies, Costa-Gomes, Crawford, and Broseta provided the definitive cognitive verification of those hierarchies, demonstrating that low-level strategic types correspond to measurable, physical operations of human information processing.
10.2 Contrast with Camerer, Ho, and Chong’s Cognitive Hierarchy Model
As the literature on bounded strategic rationality expanded, Colin Camerer, Teck-Hua Ho, and Juin-Kuan Chong introduced a prominent alternative framework in 2004: the Cognitive Hierarchy (CH) model. Both the Level-k model of Costa-Gomes et al. and the CH model formalize iterative, step-by-step strategic reasoning, but they diverge in their structural assumptions regarding how agents form beliefs about their opponents.
In the Costa-Gomes et al. Level-k specification, an agent of level $k$ maintains a degenerate belief that all other players are operating precisely at level $(k – 1)$. An L2 thinker, for example, assumes that their opponent is exclusively an L1 thinker. In contrast, the Cognitive Hierarchy model assumes that a level-$k$ agent recognizes that the population contains a mixture of all lower cognitive levels, ranging from level-0 up to level-$(k – 1)$. The distribution of these lower levels is assumed to follow a normalized Poisson distribution governed by a single mean parameter, $tau$:
$$f(m) = \frac{e^{-\tau} \tau^m}{m!}$$
This structural difference has notable implications. The Cognitive Hierarchy model is analytically parsimonious: the single parameter $tau$ describes the cognitive depth of the entire population, with empirical estimates typically converging around $\tau \approx 1.5$. Furthermore, because CH agents best-respond to a distribution rather than a single type, their predicted actions tend to be smoother and less brittle than pure Level-k predictions. However, the Costa-Gomes et al. framework offers far greater cognitive realism in initial, unlearned settings: human subjects rarely possess the mental capacity to calculate Poisson distributions of lower types, whereas the simple Level-k heuristic of best-responding directly to a single lower-order archetype mirrors the working-memory constraints revealed by MouseLab and eye-tracking process logs.
10.3 Evolution from MouseLab to High-Resolution Eye-Tracking
In the decades following the 2001 study, process-tracing methodology underwent a major technological evolution, transitioning from manual mouse-tracking software to modern, infrared corneal-reflection eye-tracking systems. While MouseLab was revolutionary for its time, it required subjects to physically steer a cursor across boxes, introducing motor friction and occasional mechanical artifacts. Eye-tracking technology eliminated this friction entirely: high-speed infrared cameras record ocular fixations, saccades, and visual regressions at frequencies exceeding 500 Hz, capturing natural gaze behavior without requiring manual mouse clicks.
Remarkably, subsequent eye-tracking replications of the Costa-Gomes, Crawford, and Broseta experiments—such as those conducted by Colin Camerer, Eric Johnson, and other behavioral researchers—overwhelmingly confirmed the core findings of the original MouseLab studies. Subjects classified as Level-1 continue to exhibit near-zero fixations on opponent payoff cells, while Level-2 players consistently display the sequential, two-phase gaze trajectory identified by the 2001 study.
Moreover, eye-tracking unlocked granular physiological dimensions that were inaccessible to MouseLab. Researchers can now measure pupillary dilation, which serves as a physiological proxy for instantaneous cognitive effort and mental workload. Modern eye-tracking studies have demonstrated that subjects executing Level-2 reasoning exhibit significant spikes in pupil dilation as they mentally compute the opponent’s anticipated choice, providing physiological confirmation of the elevated cognitive demands associated with higher-order strategic deliberation.
11. Subsequent Developments: Costa-Gomes and Crawford (2006) and Beyond
11.1 Two-Person Guessing Games and Continuous Strategy Spaces
Building upon the insights of their 2001 normal-form paper, Miguel Costa-Gomes and Vincent Crawford published a seminal follow-up study in the American Economic Review in 2006, titled “Cognition and Behavior in Two-Person Guessing Games: An Experimental Study.” While the 2001 study relied on 2×2 and 3×3 discrete normal-form games, the 2006 paper extended the investigation into large, continuous strategy spaces using two-person guessing games, conceptually related to the classic “beauty contest” games popularized by Rosemarie Nagel.
In these two-person guessing games, each player must choose a number within a continuous interval (e.g., between 100 and 900). Each player’s target score is a specified multiplier times the other player’s chosen number (e.g., Player 1’s target is $0.7 \times \text{Guess}_2$, while Player 2’s target is $1.5 \times \text{Guess}_1$). The continuous strategy space provided a critical empirical advantage: it eliminated the accidental type mimicry that occasionally plagued discrete 3×3 matrices. In a continuous guessing game, the exact number selected by an agent can identify their cognitive type with exceptional mathematical precision:
- Level-1 Guess: Best-responds to a uniform prior over the interval $[100, 900]$ (midpoint 500), selecting $0.7 \times 500 = 350$.
- Level-2 Guess: Calculates the opponent’s L1 guess and best-responds to it, choosing $0.7 \times (1.5 \times 500) = 525$.
- Nash Equilibrium Guess: Solves the simultaneous linear equations, yielding the unique mathematical fixed point.
Across sixteen continuous games with varying targets and intervals, Costa-Gomes and Crawford demonstrated that subject choices clustered around the exact numerical predictions of Level-1 and Level-2 types, completely absent choice overlap. This confirmed the predominance of shallow strategic hierarchies purely through revealed choices, reinforcing the procedural conclusions first uncovered via MouseLab search data in their 2001 study.
11.2 Applications to Strategic Communication and Deception
The theoretical insights forged in the 2001 study quickly expanded into models of asymmetric information, signaling, and strategic communication. Vincent Crawford and his coauthors applied the Level-k framework to sender-receiver games and cheap-talk environments, areas where standard equilibrium theory famously suffers from severe predictive failures. Under classical Nash formulations, cheap-talk messages in zero-sum or conflictual environments are completely uninformative: rational receivers recognize that senders have an incentive to lie, rendering all non-binding messages meaningless “babbling.”
In empirical settings, however, human communication is rife with both strategic deception and systematic credulity. By importing structural Level-k types into communication games, Crawford demonstrated that these real-world dynamics can be modeled with analytical rigor:
- Level-0 Sender: Truthfully reveals their private information or states their true intentions without strategic calculation.
- Level-1 Receiver: Credulously believes the sender’s message at face value, assuming the sender is an L0 truth-teller.
- Level-2 Sender: Anticipates the credulous Level-1 receiver and engages in active, systematic deception to manipulate the receiver’s choice.
- Level-3 Receiver: Anticipates the deceptive Level-2 sender, correctly inverting the deceptive message to identify the truth.
This structural approach to communication games illuminated the cognitive foundations of financial disclosure regulations, political discourse, and tactical negotiations. Process-tracing investigations of these communication games revealed that receivers often focus entirely on the sender’s message while ignoring the sender’s underlying incentive to mislead—a procedural signature that precisely characterizes the credulous Level-1 receiver archetype.
11.3 Integration into Mechanism Design and Applied Game Theory
The empirical confirmation that real populations consist of boundedly rational Level-k agents has profoundly impacted applied game theory and mechanism design. Standard economic mechanisms—such as the Vickrey-Clarke-Groves (VCG) auction or classical double auctions—were engineered under the assumption of universal hyper-rationality and common knowledge of equilibrium play. A celebrated theoretical pillar of this literature is the Revenue Equivalence Theorem, which states that under standard assumptions, first-price and second-price auctions yield identical expected revenues to the seller.
When evaluated under structural Level-k bidder populations, however, revenue equivalence breaks down entirely:
- In a second-price sealed-bid auction, truthful bidding is a weakly dominant strategy, making it accessible even to naive or low-level strategic thinkers who ignore their opponents’ payoffs.
- In a first-price sealed-bid auction, optimal bidding requires shading one’s bid below one’s true valuation, a calculation that requires forming beliefs about opponents’ bidding distributions. Level-1 and Level-2 bidders systematically miscalculate this bid shading, leading to empirical over-bidding relative to the risk-neutral Nash equilibrium.
Recognizing these cognitive bounds, mechanism designers now engineer non-equilibrium-robust mechanisms. By tailoring economic institutions, auction rules, and matching protocols to the empirical reality of Level-1 and Level-2 participants, economists can design market environments that prevent systematic exploitation, reduce allocative failures, and remain robust against the bounded mentalizing capacities of real human agents.
12. Enduring Legacy and the Cognitive Architecture of Strategic Thought
12.1 Epistemological Shift in Understanding Rationality in Games
The lasting legacy of Costa-Gomes, Crawford, and Broseta’s 2001 investigation lies in the fundamental epistemological shift it catalyzed across the economic discipline. By replacing the monolithic, hyper-rational actor of classical game theory with a structured, empirically validated cognitive architecture, their work redefined how economists conceptualize human agency. Prior to their paper, deviations from Nash equilibrium were routinely dismissed as transient noise, experimental anomalies, or the temporary consequences of subject confusion that would swiftly evaporate with learning.
The authors demonstrated that bounded strategic rationality is not a collection of random errors, but a deeply structured, stable, and predictable cognitive phenomenon. The human mind possesses finite working-memory buffers, constrained mentalizing capacity, and structural limits on recursive logic. These bounds are not personal failures; they are fundamental features of human biological cognition. Consequently, game theory transitioned from a purely normative, prescriptive mathematics of ideal rationality into a descriptive, empirically grounded cognitive science. Today, structural Level-k modeling is an indispensable tool across institutional design, central banking policy evaluation, antitrust analysis, and strategic management, providing a framework for forecasting human behavior where textbook equilibrium assumptions fail.
12.2 Neuroeconomic and Modern Physiological Horizons
In recent years, the behavioral hypotheses pioneered by Costa-Gomes, Crawford, and Broseta have received direct physical corroboration from neuroeconomics and cognitive neuroscience. Using functional Magnetic Resonance Imaging (fMRI), neuroscientists have mapped the neural correlates of strategic thinking as human subjects evaluate normal-form matrices. These neuroimaging studies reveal that different strategic types engage distinct anatomical brain networks:
- Level-1 Deliberation: Activates brain regions primarily associated with basic reward processing, arithmetic calculation, and self-referential cognition, such as the ventral striatum and the ventromedial prefrontal cortex (vmPFC).
- Level-2 and Higher-Order Thinking: Recruits the “Theory of Mind” (mentalizing) network, characterized by strong blood-oxygen-level-dependent (BOLD) signal activation in the medial prefrontal cortex (mPFC), the temporoparietal junction (TPJ), and the superior temporal sulcus.
When an L2 subject shifts their gaze to inspect the opponent’s payoff matrix, fMRI data confirms a simultaneous activation spike within the temporoparietal junction, the neural region responsible for representing the beliefs, desires, and intentions of external agents. By linking high-resolution process-tracing to structural econometrics and modern neurobiology, contemporary science has confirmed what Costa-Gomes, Crawford, and Broseta first demonstrated using simple cursor movements: human strategic thought is an embodied, procedural computation with distinct physiological footprints.
12.3 Synthesis: The Permanent Contribution of Costa-Gomes, Crawford, and Broseta
At the turn of the millennium, Miguel Costa-Gomes, Vincent Crawford, and Bruno Broseta set out to answer a deceptively simple question: when human beings sit down to play a game, how do they actually think? Through a brilliant convergence of experimental design, non-equilibrium decision theory, structural econometrics, and computational information tracking, they delivered one of the most comprehensive empirical investigations in the history of game theory.
Their contributions remain foundational across multiple dimensions:
- The 18-Game Panel: Their carefully calibrated sequence of asymmetric, normal-form games remains a primary empirical benchmark for evaluating non-equilibrium behavioral models.
- The Methodological Bridge: They proved that intermediate process data is an indispensable tool for resolving identification dilemmas that choice data alone can never solve.
- The Empirical Reality: They demonstrated that human strategic populations are predominantly composed of Level-1 and Level-2 thinkers, while classical Nash solvers are an extreme empirical rarity.
Ultimately, the “thinking experiments” designed by Costa-Gomes, Crawford, and Broseta fundamentally reshaped 21st-century strategic analysis. By forcing economists to confront the real cognitive mechanics of decision-makers, their work forever dismantled the fiction of the hyper-rational black box, establishing an enduring empirical architecture for understanding the complex, bounded, and deeply human nature of strategic thought.
Conclusion
The journey from the classical axioms of von Neumann, Morgenstern, and Nash to modern behavioral game theory represents one of the most significant intellectual shifts in economic thought. For decades, the hyper-rational equilibrium paradigm held sway not because it provided an accurate description of human mental operations, but because it offered a mathematically pristine architecture unburdened by psychological messiness. Yet as laboratory evidence accumulated, the gap between theoretical elegance and behavioral reality grew impossible to ignore. It was the pioneering vision of Miguel Costa-Gomes, Vincent Crawford, and Bruno Broseta that finally bridged this divide, demonstrating that cognitive bounds are not a barrier to mathematical modeling, but the very foundation upon which realistic economic theory must be built.
By leveraging the MouseLab interface to track the sequential, second-by-second information searches of human decision-makers, Costa-Gomes, Crawford, and Broseta exposed the profound empirical limitations of revealed-preference orthodoxy. Their findings proved that selecting an equilibrium action is not synonymous with executing equilibrium thought, and that human populations are fundamentally defined by shallow, bounded cognitive hierarchies. In doing so, they provided a durable blueprint for modern economics: a methodology that unites structural econometrics, psychological realism, and direct empirical observation of the decision-making process. The enduring legacy of their work is a game theory that no longer demands agents be infinitely rational, but instead celebrates the disciplined, empirical science of understanding how real human minds navigate the strategic world.
References
- Camerer, C. F. (2003). Behavioral Game Theory: Experiments in Strategic Interaction. Princeton University Press. https://press.princeton.edu/books/hardcover/9780691090061/behavioral-game-theory
- Camerer, C. F., Ho, T. H., & Chong, J. K. (2004). A cognitive hierarchy model of games. The Quarterly Journal of Economics, 119(3), 861–898. https://doi.org/10.1162/0033553041502183
- Costa-Gomes, M., Crawford, V. P., & Broseta, B. (2001). Cognition and behavior in normal-form games: An experimental study. Econometrica, 69(5), 1193–1235. https://doi.org/10.1111/1468-0262.00239
- Costa-Gomes, M. A., & Crawford, V. P. (2006). Cognition and behavior in two-person guessing games: An experimental study. American Economic Review, 96(5), 1737–1768. https://doi.org/10.1257/aer.96.5.1737
- Crawford, V. P. (2003). Lying for strategic advantage: Rational and boundedly rational misrepresentation of intentions. American Economic Review, 93(1), 133–149. https://doi.org/10.1257/000282803321455204
- Friedman, M. (1953). The methodology of positive economics. In Essays in Positive Economics (pp. 3–43). University of Chicago Press.
- Johnson, E. J., Camerer, C., Sen, S., & Rymon, T. (2002). Detecting failures of backward induction using the cognitive tool of information search. Journal of Economic Theory, 104(1), 16–47. https://doi.org/10.1006/jeth.2001.2850
- Nagel, R. (1995). Unraveling in guessing games: An experimental study. American Economic Review, 85(5), 1313–1326. https://www.jstor.org/stable/2950991
- Nash, J. (1950). Equilibrium points in n-person games. Proceedings of the National Academy of Sciences, 36(1), 48–49. https://doi.org/10.1073/pnas.36.1.48
- Simon, H. A. (1955). A behavioral model of rational choice. The Quarterly Journal of Economics, 69(1), 99–118. https://doi.org/10.2307/1884852
- Stahl, D. O., & Wilson, P. W. (1994). Experimental evidence on players’ models of other players. Journal of Economic Behavior & Organization, 25(3), 309–327. https://doi.org/10.1016/0167-2681(94)90103-1
- Stahl, D. O., & Wilson, P. W. (1995). On players’ models of other players: Theory and experimental evidence. Games and Economic Behavior, 10(1), 218–254. https://doi.org/10.1006/game.1995.1031