In 1960, British cognitive psychologist Peter Cathcart Wason published a deceptively simple experimental paper in the Quarterly Journal of Experimental Psychology titled “On the Failure to Eliminate Hypotheses in a Conceptual Task.” At a historical moment when behavioral psychology was being vigorously challenged by the nascent cognitive revolution, Wason introduced an empirical paradigm that would fundamentally dismantle naive assumptions regarding human rationality, scientific deduction, and inductive logic. The task, now universally celebrated as the Wason 2-4-6 hypothesis-testing task, confronted participants with a seemingly trivial problem: discover a latent numerical rule governing triples of numbers, seeded initially with the sequence “2, 4, 6.” What Wason uncovered was neither an issue of computational deficiency nor a transient lapse in mathematical computation. Instead, he observed a pervasive, systemic cognitive inertia: human thinkers exhibited an overwhelming propensity to confirm their initial conjectures rather than attempt their falsification.
The 2-4-6 task arrived at an intellectual crossroads dominated on one side by Karl Popper’s prescriptive philosophy of science, which demanded relentless falsificationism as the hallmark of rational inquiry, and on the other by Jean Piaget’s developmental constructivism, which posited that adult humans naturally attain a stage of formal operations characterized by hypothetico-deductive reasoning. Wason’s empirical data demonstrated that when free to generate their own test instances, university undergraduates routinely fell into a self-reinforcing epistemic trap. They formulated narrow, highly specific candidate rules—such as “numbers increasing by two” or “consecutive even numbers”—and methodically generated sequences that conformed strictly to their internalized assumptions. Because the experimenter’s latent rule was vastly broader—”any ascending sequence”—every confirming test returned a positive response, instilling a subjective, mathematically unfounded conviction of truth until the participant prematurely declared an erroneous rule.
This article provides an exhaustive, multidisciplinary analysis of the Wason 2-4-6 hypothesis-testing model. Over the course of twelve comprehensive sections, we will trace its historical and philosophical foundations, dissect its structural mechanics, examine the pivotal cognitive reinterpretation introduced by Joshua Klayman and Young-Won Ha, map the set-theoretic topologies of hypothesis spaces, review seminal experimental variations, and explore computational, neurobiological, clinical, and pedagogical ramifications. Through this investigation, the 2-4-6 paradigm emerges not merely as a historical laboratory curiosity, but as a foundational bedrock for understanding the cognitive architecture of human belief formation, information search, and epistemic fallibility.
1. Historical Foundations and the Genesis of the 2-4-6 Paradigm
The emergence of the 2-4-6 task cannot be understood in isolation from the broader epistemological and methodological shifts that transformed psychological science in the mid-twentieth century. As experimental psychology emancipated itself from the mechanistic constraints of stimulus-response behaviorism, researchers increasingly sought empirical frameworks capable of interrogating the active, generative nature of human thought.
1.1 Peter Cathcart Wason and the Cognitive Revolution
During the late 1950s, experimental psychology underwent an epochal transformation commonly referred to as the cognitive revolution. For decades, radical behaviorism, championed by figures like B. F. Skinner, had relegated internal mental operations to an unobservable “black box,” asserting that human action could be sufficiently explained through schedules of reinforcement and conditioned reflexes. However, the explanatory inadequacies of behaviorism when confronted with complex language acquisition, problem-solving, and abstract reasoning catalyzed a return to mentalistic concepts. Peter Cathcart Wason, working within the intellectually vibrant milieu of University College London (UCL), found himself heavily influenced by Sir Frederic Bartlett’s cognitive approaches to remembering and thinking, as well as the pioneering work of Jerome Bruner, Jacqueline Goodnow, and George Austin in their landmark 1956 volume, A Study of Thinking.
Bruner and his colleagues had demonstrated that human concept attainment was an active, strategic process involving hypothesis testing. Yet Wason perceived a profound limitation in their experimental protocols: concept attainment studies typically utilized reception paradigms or constrained selection paradigms, where subjects were presented with pre-fabricated cards featuring geometric designs and asked to categorize them. Wason recognized that such fixed-choice psychometric designs failed to capture the open-ended, self-directed nature of real-world scientific inquiry. In actual scientific practice, researchers are not passive recipients of pre-packaged stimuli; they must actively formulate candidate hypotheses and invent novel experiments to interrogate nature. Driven by an abiding interest in the systematic anomalies and fallacies of human reasoning, Wason sought to construct an open-ended discovery task that granted participants complete structural freedom to generate their own empirical data points, thereby unveiling the spontaneous strategies—and systemic pathologies—of the human mind in its natural investigative state.
1.2 Popperian Epistemology and Falsificationism
Wason’s methodological architecture was directly inspired by the philosophical tenets of Karl Popper, whose 1935 magnum opus, Logik der Forschung, was translated into English as The Logic of Scientific Discovery in 1959. Popper mounted a radical assault against logical positivism and naive inductivism. The inductivist paradigm posited that universal scientific laws are inductively derived and incrementally confirmed through the steady accumulation of positive empirical observations. Popper exposed the fatal logical asymmetry at the heart of this doctrine: no finite quantity of confirming observations can ever definitively establish the truth of a universal statement of the form “All As are Bs,” because future observations may always contradict it. Conversely, a single authentic counter-instance possessing the properties “A and not-B” deductively refutes the universal claim via the valid logical inference of modus tollens.
Consequently, Popper proposed falsifiability as the primary demarcation criterion separating genuine empirical science from pseudo-science. According to Popperian epistemology, a rational investigator does not design experiments to verify a favored hypothesis; rather, the scientist relentlessly formulates risky, bold conjectures and immediately attempts to refute them through severe, disconfirming empirical tests. Theories that survive rigorous attempts at falsification achieve provisional corroboration, but they remain forever tentative hypotheses awaiting potential refutation. Wason seized upon this philosophical mandate and transformed it into an empirical psychological benchmark. He asked a fundamentally novel question: Do human beings, when functioning as miniature scientists in a controlled inductive problem-solving environment, intuitively adopt a Popperian hypothetico-deductive posture, or do they default to naive inductivism and illicit verification?
1.3 The Seminal 1960 Experiment
In his 1960 paper, Wason presented the initial empirical operationalization of his hypothesis-testing paradigm. The experimental design was radically minimalist. Wason recruited 29 university students (a cohort of intellectually privileged individuals presumed to possess superior logical aptitude) and introduced them to a numerical discovery game. The experimenter informed the participants that he had in mind a specific rule governing sequences of three numbers (triples). To initiate the task, the experimenter presented a single positive exemplar: the triple 2-4-6. The participants were instructed that their objective was to discover the latent rule by generating their own triples of three numbers. After each triple was presented, the experimenter would deliver immediate categorical feedback, stating strictly whether the sequence conformed or did not conform to the secret rule.
Crucially, participants were required to record not only the triple they wished to test, but also the underlying hypothesis or reason that motivated the selection of that specific triple. They were allowed to generate as many triples as they deemed necessary, but they were instructed to announce their definitive rule only when they were highly confident that they had discovered the correct formulation. The latent rule conceived by Wason was deliberately and deceptively general: “any three numbers in increasing order of magnitude” (or simply ascending numbers). The initial seed triple, 2-4-6, was an intentional cognitive trap; it possessed a multitude of salient, mathematically appealing properties: even numbers, an arithmetic progression with a common difference of two ($x, x+2, x+4$), and a continuous proportional relationship.
The quantitative results shocked Wason and sent shockwaves through experimental psychology. Out of the 29 intelligent university students, only 6 participants (20.7%) discovered the correct rule on their first announced attempt without articulating an incorrect rule beforehand. Even after being told that their initial declared rules were wrong and instructed to continue testing, only a minority successfully converged on the broad, correct rule without profound cognitive struggle. The overwhelming majority of participants immediately generated hypotheses of extreme specificity (e.g., “intervals of two,” “ascending even numbers,” “the second number is the average of the first and third”). To test these ideas, they systematically generated triples that perfectly matched their hyper-specific hypotheses (e.g., testing 8-10-12, 14-16-18, 20-22-24). Every single one of these triples conformed to Wason’s broad ascending rule, eliciting consistent “Yes” feedback from the experimenter. This constant stream of positive feedback produced an intense, subjective illusion of validity, leading participants to formally announce their narrow hypotheses with absolute conviction, only to be bewildered when informed that their assertions were fundamentally incorrect.
2. Structural Mechanics and Experimental Protocol of the 2-4-6 Task
To appreciate the psychological dynamics of the 2-4-6 task, one must examine its mechanical, formal, and methodological structure. Unlike deductive tasks where all premises are explicitly supplied, the 2-4-6 paradigm operates within an underdetermined, open-ended problem space that requires simultaneous information generation, hypothesis formulation, and metacognitive monitoring.
2.1 Standard Task Administration and Rules
The standard administrative protocol developed by Wason, and subsequently standardized across hundreds of psychological laboratories, begins with precise, standardized verbal and written instructions. The subject is seated opposite the experimenter and provided with a formal response sheet divided into sequential columns: one for the proposed triple of numbers, one for the subject’s explicit rationale or current working hypothesis, and one reserved for the experimenter’s binary feedback. The experimenter states clearly:
“I have in mind a rule that relates sets of three numbers. The triple 2-4-6 conforms to this rule. I would like you to discover what this rule is. You may do this by writing down sets of three numbers on this sheet, together with the reason why you chose them. I will then tell you whether your numbers conform to my rule or not. You may test as many sets of numbers as you like. When you are confident that you have discovered the rule, and only then, you are to tell me what the rule is.”
The administration enforces several strict experimental constraints. The feedback delivered by the experimenter is strictly dichotomous, verbalized exclusively as “conforms” (positive feedback) or “does not conform” (negative feedback), without any elaborative qualitative guidance or tonal inflection. If a participant presents an invalid sequence or an ambiguously written hypothesis, the experimenter requests clarification without offering strategic hints. Furthermore, the termination criteria are absolute: participants are explicitly discouraged from guessing. They are urged to continue generating data until they achieve subjective certainty. When a participant formally announces a rule, if it does not match the experimenter’s latent specification, the experimenter unequivocally rejects the declaration, stating: “That is not the rule I have in mind. Please continue testing.” This cycle repeats until the participant successfully identifies the latent rule or reaches an experimental abandonment threshold (typically capped at 45 to 60 minutes or a predefined maximum number of trials).
2.2 The Latent Mathematical Rule: Any Ascending Sequence
The cognitive genius of the 2-4-6 task resides in the profound mathematical and structural asymmetry between the latent target rule ($T$) and the initial candidate hypotheses ($H$) naturally evoked by the seed exemplar. The true target rule chosen by Wason is:
$T = {(x, y, z) in mathbb{R}^3 mid x < y < z}$
In set-theoretic terms, $T$ represents the expansive set of all real (or integer) ordered triples exhibiting strictly monotonic positive progression. This rule is remarkably simple; it requires no arithmetic parity, no fixed intervals, and no complex algebraic dependencies. However, the seed triple presented to the participant, $(2, 4, 6)$, exhibits an exceptionally dense concentration of salient mathematical properties. It is simultaneously:
- An ascending arithmetic progression with a common difference of two ($x, x+d, x+2d$ where $d=2$).
- A sequence of consecutive positive even integers.
- A linear sequence governed by the functional form $f(n) = 2n$.
- A set where the middle term is the exact arithmetic mean of the extremes ($y = \frac{x+z}{2}$).
- A sequence where the third term is the summation of the first two ($z = x + y$).
Because the seed triple possesses all these nested structural characteristics, human pattern-recognition mechanisms instantly fire, generating highly restrictive hypotheses. A typical participant immediately assumes the rule is “numbers increasing by two” ($H_1$), “even numbers increasing by two” ($H_2$), or “arithmetic sequences” ($H_3$). Structurally, these hypothesized sets represent proper, embedded subsets of the true rule: $H_1 \subset T$, $H_2 \subset T$, and $H_3 \subset T$. This embedded topology constitutes what cognitive scientists refer to as the hierarchical generality problem. Any sequence generated to satisfy $H_1$ (such as 10-12-14 or 100-102-104) is inherently an ascending sequence and therefore unconditionally satisfies $T$. Consequently, the participant receives nothing but reinforcing positive feedback, remaining blissfully blind to the vast, encompassing expanse of $T$ that lies beyond the self-constructed perimeter of $H$.
2.3 Data Collection Metrics in Empirical Replications
Over decades of empirical replication, cognitive psychologists have operationalized a rich taxonomy of quantitative and qualitative metrics to rigorously quantify participant performance in the 2-4-6 paradigm:
- Trial Sequence Length (TSL): The total number of unique triples generated by a participant prior to their first formal rule announcement, as well as the aggregate number of trials required to reach ultimate task resolution.
- Hypothesis-Test Congruence (Instance Classification): Every generated triple is classified in relation to the participant’s stated working hypothesis. A test is defined as a positive test ($+test$) if the triple is an exemplar that the participant’s hypothesis predicts should conform to the rule. A test is defined as a negative test ($-test$) if the triple is an instance that the participant’s hypothesis predicts should fail to conform to the rule.
- Disconfirmation Latency: The precise temporal duration and the number of intervening trials between the reception of negative feedback and the generation of an alternative conceptual hypothesis.
- Rule Announcement Accuracy: The proportion of participants achieving task success on the First Hypothesis Announcement (FHA), versus those requiring multiple announcements (Second or Subsequent Hypothesis Announcements), versus those who terminate in complete failure or cognitive surrender.
- Concurrent Think-Aloud Protocols: Utilizing the qualitative methodology formalized by K. Anders Ericsson and Herbert Simon, researchers capture continuous verbal protocols during testing. These recordings are transcribed and coded to track how participants update their internal mental models, register surprise, and rationalize unexpected feedback.
3. Cognitive Architecture: Verification Bias versus Falsification Failure
The persistent failure of participants in the original 2-4-6 task compelled Wason to construct a psychological explanation rooted in human cognitive architecture. He initially diagnosed this phenomenon as an intrinsic “verification bias”—a deep-seated, irrational cognitive pathology wherein the human mind systematically avoids negative evidence and actively seeks to confirm its pre-existing mental representations.
3.1 The Operational Definition of Verification Bias
In Wason’s classic conceptual framework, verification bias denotes the active, unidirectional search for evidence that corroborates an established or tentatively adopted proposition. It manifests as a profound psychological resistance to generating any experimental test that could potentially disconfirm, destabilize, or invalidate the reigning hypothesis. When placed in the 2-4-6 task, participants do not behave as neutral, dispassionate information processors. Instead, once an initial conjecture—such as “the interval between numbers must be equal”—crystallizes in their working memory, they treat that conjecture as an established truth to be proven rather than an unverified possibility to be interrogated.
This bias is not simply a passive failure to consider alternatives; it is an active, selective behavioral orientation. Participants deliberately construct test instances that represent the purest, most stereotypical exemplars of their internal hypothesis. When an individual operating under the hypothesis “numbers increasing by two” generates the triple 8-10-12, they expect, desire, and anticipate positive feedback. When the experimenter affirms “that conforms,” the participant experiences an immediate reduction in epistemic ambiguity. This creates a self-reinforcing positive feedback loop: each confirming instance artificially elevates subjective confidence while providing zero diagnostic informational value. The participant becomes trapped within a subjective illusion of certainty, confusing the mere accumulation of positive instances with genuine scientific proof, entirely oblivious to the fact that their tests have completely failed to probe the boundary conditions of their hypothesis.
3.2 The Asymmetry of Information in Rule Search
To understand why Wason characterized this behavior as fundamentally irrational, one must appreciate the stark informational asymmetry inherent in hypothesis spaces. In any formal inductive search where the hypothesized rule $H$ is a strict subset of the latent rule $T$ ($H subset T$), positive feedback possesses precisely zero bits of diagnostic information for discriminating between $H$ and $T$. If a participant hypothesizes “even numbers increasing by two” and tests 20-22-24, receiving the feedback “conforms,” this outcome is mathematically predicted by their hypothesis ($H$). However, it is also equally predicted by the hypothesis “any three even numbers,” “any numbers increasing by two,” “any ascending numbers,” and “any three numbers whatsoever.” Therefore, positive feedback fails to eliminate a single competing hypothesis from the nested hierarchy.
Conversely, negative feedback carries immense informational utility. If a participant hypothesizes “numbers increasing by two” and purposefully generates a sequence that violates this rule—such as 2-4-7 (an interval of 2 followed by 3) or 5-4-3 (a descending sequence)—the outcome yields profound diagnostic clarity:
- If 2-4-7 receives the feedback “conforms”, the hypothesis “intervals of two” is instantly, definitively, and irrevocably falsified via modus tollens. The participant is instantly liberated from an erroneous mental model and compelled to broaden their conceptual horizon.
- If 2-4-7 receives the feedback “does not conform”, the participant has successfully identified an empirical boundary condition, establishing that unconstrained variability in step size is impermissible.
Despite this clear informational superiority, human participants exhibit an overwhelming cognitive aversion to negative feedback. Processing disconfirming evidence imposes a massive cognitive load: it forces working memory to dismantle an existing mental model, resolve acute cognitive dissonance, and initiate an entirely new conceptual search across an infinite mathematical problem space. Positive evidence, by contrast, requires minimal cognitive effort, as it seamlessly integrates into pre-existing schemas, sustaining epistemic comfort at the expense of empirical accuracy.
3.3 Wason’s Original Diagnostic Classification
In analyzing the trial-by-trial behavior of his participants, Wason categorized the observed methodologies into three distinct behavioral archetypes: genuine confirmation, pseudo-falsification, and genuine falsification.
Genuine Confirmation characterized the vast majority of sequences. Here, participants exclusively generated triples that conformed to their active hypothesis, predicting positive outcomes and seeking direct verification. If their hypothesis was “multiples of three,” they tested 3-6-9, 9-12-15, and 30-33-36, completely refusing to step outside the boundaries of their assumed rule.
Pseudo-Falsification represented a fascinating, subtle cognitive aberration identified by Wason. In these instances, a participant appeared to generate a disconfirming triple, but upon closer examination, the test was designed in such a way that it could not genuinely threaten their core hypothesis. A participant might adopt the hypothesis “numbers increasing by two” and test the triple 1-3-5. When asked for their rationale, they might state: “I wanted to see if it works with odd numbers.” Crucially, this is not a falsification test of the structural rule “increasing by two”; it is merely the positive testing of a secondary, peripheral variation. True pseudo-falsification occurs when a subject generates a triple that they fully anticipate will receive a “does not conform” response, but which does not structurally isolate the critical causal variable. If a subject hypothesizing “consecutive even numbers” tests 2-4-5, predicting failure, and receives “conforms,” they are frequently unable to comprehend the result, often dismissing it as an anomaly or redefining the rule post-hoc to include bizarre ad-hoc exemptions.
Genuine Falsification was exceedingly rare. It occurred only when a participant deliberately formulated a triple that violated their primary conceptual conjecture for the explicit purpose of testing whether the rule held in the absence of that property. For example, an individual who deeply believed that the rule required equal arithmetic steps would intentionally test 1-2-10 specifically to determine if the equality of intervals was a mandatory structural constraint. Wason concluded that the human cognitive apparatus is naturally deficient in the capacity for genuine falsification. He argued that when confronted with open-ended inductive challenges, human reasoning is governed not by formal logic or rational deduction, but by a systematic, uncritical drive toward empirical self-justification—a conclusion that fundamentally challenged optimistic assessments of human intellectual rationality.
4. The Klayman and Ha Reformulation: Positive Test Strategy (PTS)
For more than a quarter of a century, Wason’s interpretation of the 2-4-6 task stood as an unassailable pillar of cognitive psychology, cited extensively as definitive empirical proof of irrational human confirmation bias. However, in 1987, Joshua Klayman and Young-Won Ha published a revolutionary theoretical critique in Psychological Review that fundamentally transformed the scientific understanding of the 2-4-6 task and dismantled Wason’s diagnostic claims.
4.1 Deconstructing Confirmation Bias (1987)
Klayman and Ha argued that Wason had committed a fundamental conceptual error: he had conflated a specific cognitive strategy for selecting test instances with a psychological motivation to confirm pre-existing beliefs. In their seminal paper, “Confirmation Bias Re-examined,” they demonstrated that what Wason had labeled “verification bias” was, in reality, the execution of a general-purpose, content-neutral heuristic which they termed the Positive Test Strategy (PTS).
Klayman and Ha rigorously separated the logical structure of a test from the psychological intent of the tester. They defined the two operational testing strategies as follows:
- Positive Test Strategy ($+test$): Testing an instance that is hypothesized to possess the target property, or an instance that falls directly within the boundaries of the hypothesized rule ($x in H$).
- Negative Test Strategy ($-test$): Testing an instance that is hypothesized not to possess the target property, or an instance that falls entirely outside the boundaries of the hypothesized rule ($x notin H$).
Wason had assumed that a $+test$ is synonymous with an intention to verify, and that a $-test$ is synonymous with an intention to falsify. Klayman and Ha proved mathematically that this equivalence is fundamentally false. A positive test strategy can, under appropriate structural conditions, lead directly to catastrophic empirical falsification. Conversely, a negative test strategy can easily result in verification. Crucially, when an individual applies a positive test strategy, they are simply asking: “Let me examine an instance where I expect the phenomenon to occur, to see if it actually occurs.” This is not a pathological refusal to accept counter-evidence; it is an intuitive heuristic for probing hypothesis spaces based on expected presence rather than expected absence.
4.2 PTS as an Adaptive Heuristic
Why would the human cognitive apparatus default to a Positive Test Strategy? Klayman and Ha answered this question through the lens of ecological rationality and evolutionary adaptation, prefiguring the bounded rationality paradigms of Herbert Simon and Gerd Gigerenzer. In the real, natural world, the properties and phenomena that organisms must learn about are not distributed symmetrically across the environment. Real-world target phenomena are overwhelmingly rare events with extremely low base rates.
Consider a physician attempting to diagnose an exceptionally rare medical disease, or an early human attempting to discover which species of mushrooms are fatally poisonous. In such sparse informational ecologies, the set of all entities that possess the target property ($T$) is microscopic compared to the infinite set of entities that do not possess it ($\bar{T}$). If a researcher wishes to test the hypothesis that “ingesting Amanita phalloides causes hepatic necrosis,” it is profoundly uninformative to execute a negative test strategy by having a patient consume a strawberry, an apple, a rock, or a glass of water to see if hepatic necrosis fails to develop. In a world of sparse phenomena, testing non-cases ($-tests$) yields virtually infinite uninformative verifications; almost everything in the universe does not cause hepatic necrosis.
Under realistic conditions where both the hypothesized rule $H$ and the true rule $T$ have low base rates in the universal domain, the Positive Test Strategy is mathematically optimal. When a researcher tests an instance where the property is predicted to be present ($x in H$), the test has an extraordinarily high probability of yielding a decisive, falsifying outcome if the hypothesis is incorrect—specifically, a false positive (the experimenter discovers that the predicted phenomenon fails to occur). Thus, in naturalistic environments, PTS functions as a brilliant, highly efficient, and adaptive cognitive heuristic that maximizes the probability of rapid error detection with minimal computational overhead.
4.3 Implications for the Reinterpretation of Wason’s Findings
The profound revelation of Klayman and Ha’s analysis is that Peter Wason’s 2-4-6 task does not represent the real world. Instead, the 2-4-6 paradigm represents an atypical, highly artificial, and mathematically perverse informational environment. The 2-4-6 task is pathological, not human cognition.
In the 2-4-6 task, the target rule $T$ (“any ascending sequence”) is extraordinarily broad; it encompasses a massive proportion of all possible numerical triples. Meanwhile, the initial hypotheses $H$ formed by human participants (“increasing by two,” “even numbers”) are extraordinarily narrow. Thus, the 2-4-6 task forces the relationship between the hypothesis and the truth into a catastrophic topological alignment: the embedded subset ($H subset T$).
In an embedded subset topology, the Positive Test Strategy suffers an absolute mathematical failure. Because every single member of $H$ is also a member of $T$, testing instances within $H$ can never produce a false positive. The experimenter can never say “does not conform” to an exemplar generated from $H$. The only way to falsify the hypothesis in an embedded subset configuration is to execute a negative test strategy (testing instances outside $H$). But human beings, evolutionary adapted for natural environments where PTS is highly effective, deploy their standard, highly successful heuristic. They test instances within $H$, hit the mathematical blind spot of the subset trap, and receive uniform positive reinforcement.
Therefore, Klayman and Ha completely rehabilitated human rationality in the 2-4-6 paradigm. The participants in Wason’s experiments were not irrational actors suffering from a dogmatic verification bias. Rather, they were rational organisms deploying an ecologically adaptive cognitive heuristic (PTS) within an engineered, counter-intuitive task environment specifically designed to expose that heuristic’s rare structural vulnerability. Wason’s original diagnosis of intrinsic human irrationality was successfully transformed into a classic demonstration of cognitive-environmental mismatch.
5. Set-Theoretic Topologies of Hypothesis-Rule Relationships
To establish a rigorous mathematical foundation for the dynamics of the 2-4-6 task, one must formally categorize the structural relations that can exist between the hypothesized rule set ($H$) and the true target rule set ($T$). The probability of hypothesis falsification under the Positive Test Strategy ($+test$) versus the Negative Test Strategy ($-test$) is entirely dictated by the set-theoretic topology governing the interaction of $H$ and $T$ within the universal domain of numerical triples ($U$).
5.1 The Embedded Subset Relationship
The embedded subset configuration represents the exact architectural trap engineered by Wason in the canonical 2-4-6 paradigm. Formally, this condition is defined as:
$H \subset T \quad \text{such t\hat} \quad H \neq T \quad \text{and} \quad (T setminus H) \neq \emptyset$
Under this topological condition, every element belonging to the hypothesized set $H$ is completely contained within the target set $T$. The consequences for cognitive search strategies are absolute:
- Executing a Positive Test Strategy ($x in H$): Because $H subset T$, any element $x$ selected from $H$ must necessarily be an element of $T$. Therefore, the feedback from the experimenter will unequivocally be “conforms” on 100% of trials. The probability of discovering an error via positive testing is precisely zero: $P(\text{False Positive} mid x in H) = 0$. The participant is trapped in an infinite positive feedback loop, systematically reinforcing an erroneous, overly specific belief.
- Executing a Negative Test Strategy ($x in \bar{H}$): To discover that the hypothesis $H$ is incorrect, the participant is mathematically compelled to sample an element from the complement of $H$ that simultaneously resides within $T$—that is, the set difference $x in (T setminus H)$. If the participant tests a sequence outside their hypothesis (for example, testing 1-5-10 when their hypothesis is “increasing by two”), the sequence does not belong to $H$, yet the experimenter announces “conforms.” This unexpected positive feedback immediately exposes the inadequacy of $H$, proving that the true rule is broader than the candidate hypothesis.
The embedded subset relationship is the sole topological configuration where the Positive Test Strategy is completely blind to falsification. The tragic cognitive reality of Wason’s task is that the seed triple 2-4-6 universally evokes hypotheses that are strict subsets of “ascending numbers,” thereby ensuring that the default cognitive heuristic of human reasoning inevitably fails.
5.2 The Superset Relationship
The inverse topological configuration occurs when the hypothesized rule set $H$ is broader than the true target rule set $T$. Formally, this is the superset relationship:
$T \subset H \quad \text{such t\hat} \quad T \neq H \quad \text{and} \quad (H setminus T) \neq \emptyset$
In this structural arrangement, the participant’s mental conjecture is excessively expansive, encompassing instances that the true rule does not permit. Imagine an experimental scenario where the experimenter’s latent target rule $T$ is “consecutive even numbers” (e.g., 2-4-6, 8-10-12), but the participant mistakenly adopts the broad hypothesis $H$: “any three numbers in increasing order of magnitude.”
In this superset environment, the behavioral outcomes of testing strategies are completely reversed:
- Executing a Positive Test Strategy ($x in H$): When the participant deploys PTS, they sample instances from their expansive hypothesis $H$. Inevitably, they will select an instance that resides within $(H setminus T)$—for example, testing the ascending sequence 1-3-7 or 2-5-9. When presented with these sequences, the experimenter immediately delivers negative feedback: “does not conform.” This single negative test instance constitutes a decisive false positive, immediately falsifying the over-generalized hypothesis $H$ and demonstrating to the participant that their conjecture is too broad.
- Executing a Negative Test Strategy ($x in \bar{H}$): If the participant were to test an instance outside their broad hypothesis (such as a descending sequence, 10-8-6), it falls entirely outside $T$, confirming what they already anticipated and yielding minimal diagnostic utility for refining the boundaries of $T$.
Empirical research extensively demonstrates that when human participants are structurally situated in a superset topology, their rate of hypothesis discovery accelerates exponentially. Because their natural inclination toward the Positive Test Strategy directly generates falsifying counter-examples, participants rapidly converge upon the correct rule without requiring explicit methodological training.
5.3 Overlapping and Disjoint Sets
The remaining set-theoretic configurations comprise overlapping (intersecting) sets and entirely disjoint sets, both of which demonstrate the robust utility of the Positive Test Strategy in non-embedded spaces.
Overlapping Sets (Intersection): Formally defined as:
$H \cap T \neq \emptyset, \quad \text{where} \quad (H setminus T) \neq \emptyset \quad \text{and} \quad (T setminus H) \neq \emptyset$
In an overlapping topology, the participant’s hypothesis shares a common domain with the true rule, but each set contains elements excluded from the other. For example, suppose the true rule is “all numbers must be strictly increasing” ($T$), but the participant hypothesizes “the sum of the first two numbers equals the third number” ($H$). The seed triple 2-4-6 belongs to both sets ($2+4=6$ and $2 < 4 < 6$).
If the participant deploys the Positive Test Strategy within $H$, they might generate the triple 10-20-30 ($10+20=30$). Because this sequence also happens to be ascending, it falls within $H cap T$, returning a positive “conforms” response. However, if the participant subsequently tests the triple 5-2-7 ($5+2=7$), this sequence belongs to $H$ but violates $T$ because it is not ascending. The experimenter immediately responds “does not conform.” Through this false positive, PTS successfully achieves falsification, proving to the participant that their linear summation hypothesis is unsustainable. Thus, in overlapping topologies, PTS possesses intrinsic falsifying capability, with the speed of falsification being a direct mathematical function of the relative area of $(H setminus T)$ to the total area of $H$.
Disjoint Sets: Formally defined as:
$H cap T = emptyset$
In this extreme condition, the participant’s hypothesis shares zero elements with the latent target rule. If the true rule is “ascending numbers” and the participant somehow hypothesizes “all numbers must be strictly descending,” every single positive test generated from $H$ (e.g., 9-8-7) immediately returns negative feedback (“does not conform”). Falsification is total, immediate, and instantaneous on the very first trial. The following taxonomy summarizes these topological dynamics:
- Embedded Subset ($H subset T$): PTS yields 100% positive confirmations ($0%$ falsification). Falsification is achievable exclusively via Negative Test Strategy (testing elements of $\bar{H}$). This is Wason’s specific trap.
- Superset ($T subset H$): PTS yields both confirmations and falsifying false positives. Falsification occurs naturally and rapidly through positive testing alone.
- Overlapping ($H \cap T \neq \emptyset$): PTS yields confirmations within the intersection and decisive falsifications within $(H setminus T)$. Highly diagnostic in ecologically normal environments.
- Disjoint ($H cap T = emptyset$): PTS yields instantaneous 100% falsification on Trial 1.
6. Empirical Variations and Structural Manipulations of the Paradigm
Following the recognition that Wason’s original 1960 protocol represented an artificially constrained task environment, cognitive psychologists launched a decades-long empirical program designed to manipulate the linguistic, structural, and social parameters of the 2-4-6 paradigm. These studies revealed that slight modifications to the task’s architecture could dramatically attenuate the apparent “irrationality” of human subjects.
6.1 Tweney’s Dual-Rule DAX/MED Paradigm
The most influential and theoretically profound empirical variation of the 2-4-6 task was developed in 1980 by Ryan Tweney and his colleagues at Bowling Green State University. Tweney recognized that the standard Wason protocol suffered from a profound linguistic and cognitive asymmetry: the experimenter’s feedback was inherently evaluative, casting non-conforming triples into a linguistic abyss of failure (“does not conform” or “No”). Psychologically, human beings struggle to extract constructive meaning from pure negation.
To eliminate this negative valence, Tweney, Michael Doherty, and their collaborators invented the DAX/MED paradigm. In this formulation, participants were informed that the experimenter had two distinct rules in mind: triples that conformed to the first rule were classified under the artificial label DAX, whereas triples that conformed to the second rule were classified under the label MED. The experimenter presented 2-4-6 as an exemplar of a DAX triple. Participants were then instructed to generate triples, and after each test, the experimenter would categorize the triple neutrally as either DAX or MED. In reality, the experimenter’s underlying operational logic was structurally identical to Wason’s: a triple was DAX if its numbers were in ascending order of magnitude, and MED if they were not in ascending order of magnitude.
The experimental consequences of this simple reframing were astonishing. In the standard Wason paradigm, success rates on the first hypothesis announcement hovered between 15% and 25%. In Tweney’s DAX/MED condition, task success skyrocketed to between 60% and 80%. Why did this occur?
The introduction of the complementary label MED completely transformed the cognitive utility of the Positive Test Strategy. In the canonical task, to falsify the hypothesis “increasing by two,” a participant had to generate a negative test (a sequence they expected to fail, eliciting a useless “No”). In the DAX/MED task, participants could continue to rely entirely on their preferred Positive Test Strategy, but they could now direct it toward discovering the properties of the alternative category MED. Participants would hypothesize: “Perhaps MED means the numbers are decreasing,” and they would positively test this hypothesis by generating 6-4-2. The experimenter would respond: “That is a MED.” By actively seeking positive instances of MED, participants effortlessly generated sequences that were non-ascending. This empirical exploration of the complement space rapidly unmasked the true boundary conditions of DAX, allowing participants to recognize that DAX simply meant “any ascending sequence.” Tweney’s work proved conclusively that cognitive failures in the 2-4-6 task were largely driven by semantic framing and the psychological aversion to negative verbal feedback, rather than an inability to engage in inductive reasoning.
6.2 Linguistic Framing and Conversational Implicature
A parallel line of critical inquiry arose from the domain of linguistic pragmatics, grounded in H. Paul Grice’s theory of conversational implicature. Psycholinguists pointed out that psychological experiments are, fundamentally, communicative social interactions governed by implicit conversational conventions. According to the Gricean Maxim of Quantity (be as informative as necessary) and the Maxim of Relation (be relevant), an interlocutor assumes that an information provider will provide salient, pertinent data.
When an experimenter presents an intelligent university student with the specific seed sequence 2-4-6, the student does not treat this sequence as a randomly drawn mathematical tuple from an infinite universe. Instead, applying standard pragmatic reasoning schemas, the subject assumes that the experimenter chose 2-4-6 deliberately, and that its most conspicuous perceptual and mathematical characteristics—namely, that the numbers are even and increase by equal intervals of two—are directly communicative and diagnostic of the intended rule. The presentation of 2-4-6 acts as a powerful conversational cue that misdirects the subject’s attention toward narrow arithmetic regularities.
Empirical studies that systematically manipulated the seed triple provide compelling support for this pragmatic critique. When researchers initiated the task with non-stereotypical, minimally salient ascending triples—such as 1-7-13, 2-8-9, or 103-412-999—the rate of immediate rule discovery increased significantly. In these conditions, the absence of clean arithmetic symmetries (such as consecutive even numbers) prevented the conversational implicature from taking root. Participants were immediately forced to formulate broader relational hypotheses, such as “numbers increasing in size,” completely bypassing the narrow subset traps that ensnared Wason’s original subjects.
6.3 Instructional Interventions and Counterfactual Prompting
Researchers have also extensively investigated whether explicit pedagogical prompts, instructions, or cognitive debiasing interventions can induce genuine falsification within the standard 2-4-6 framework. The empirical results of these interventions reveal deep insights into the malleability of human reasoning:
- Direct Falsification Instructions: Wason and subsequent investigators attempted to explicitly command participants to adopt a critical posture: “Try to prove that your rule is wrong,” or “Generate triples that you think do not conform to the rule.” Astonishingly, these direct commands exhibited minimal efficacy. Untrained participants frequently misunderstood how to falsify their own ideas, generating pseudo-falsifications or continuing to use positive testing while claiming they were looking for refutations. Direct exhortation to “falsify” fails because it tells participants *what* to achieve without providing the operational set-theoretic algorithms required to execute it.
- The “Consider-the-Opposite” Strategy: Developed by Charles Lord, Mark Lepper, and colleagues, this intervention requires participants to explicitly write down at least two alternative or competing hypotheses before they are permitted to generate any test triples. By forcing the cognitive architecture to entertain multiple rival mental models simultaneously, working memory is decoupled from exclusive commitment to the primary seed hypothesis. This technique dramatically elevates the proportion of negative and diagnostic tests, substantially improving first-announcement accuracy.
- Social and Collaborative Framing: Research examining collaborative hypothesis testing demonstrates that when individuals work in dyadic or triadic groups, task success increases markedly. In collaborative contexts, social dynamics naturally foster adversarial reasoning: one participant’s favored hypothesis is greeted with skepticism by another, prompting group members to propose triples designed to test competing claims. This social distribution of cognitive labor naturally simulates the peer-review mechanisms of the broader scientific community, effectively mitigating individual-level heuristic vulnerabilities.
7. Individual Differences, Cognitive Capacity, and Metacognition
While experimental manipulations demonstrate the profound impact of task architecture, performance on the 2-4-6 task also varies substantially across individuals. Modern differential psychology and cognitive neuroscience have illuminated the specific psychometric, cognitive, and metacognitive variables that predict an individual’s capacity to navigate the 2-4-6 problem space successfully.
7.1 Working Memory Capacity and Executive Function
The execution of rational hypothesis testing places extraordinary demands on the human central executive. According to the multi-component model of working memory developed by Alan Baddeley, solving the 2-4-6 task without falling into the confirmation trap requires the continuous maintenance, manipulation, and updating of abstract information under severe cognitive load.
To successfully uncover the latent rule, a participant must:
- Maintain the active candidate hypothesis in focal attention.
- Simultaneously represent at least one alternative, mutually exclusive hypothesis.
- Construct a novel sequence of numbers that violates the candidate hypothesis while satisfying the alternative.
- Inhibit the potent, automatic impulse to test the primary hypothesis (prefrontal inhibitory control).
- Integrate the experimenter’s feedback, correctly updating the probability distribution across the hypothesis space.
Empirical studies measuring Working Memory Capacity (WMC)—utilizing standardized psychometric instruments such as the Automated Operation Span Task (O-Span) or the Reading Span Task—demonstrate a robust, statistically significant positive correlation between WMC and performance on the 2-4-6 task. Individuals with high working memory capacity are vastly more likely to generate alternative candidate hypotheses prior to announcing their rules. Conversely, individuals with restricted working memory resources experience cognitive overload. Unable to sustain multiple conceptual models in working memory, their cognitive apparatus collapses into attentional tunneling: they anchor obsessively to their initial salient pattern (e.g., “+2”), generating perseverative, confirmatory sequences in an effort to minimize mental strain.
7.2 Cognitive Reflection and Need for Cognition
Beyond raw computational capacity, individual performance on the 2-4-6 task is profoundly governed by cognitive style and dispositional reflectivity. Within the framework of Dual-Process Theory—formalized by cognitive scientists such as Jonathan Evans, Keith Stanovich, and Daniel Kahneman—human cognition is mediated by two distinct modes of information processing:
- Type 1 Processing (Heuristic/Intuitive): Autonomous, rapid, cognitively effortless, and heavily reliant on associative pattern recognition. When presented with 2-4-6, Type 1 processing instantly generates intuitive arithmetic regularities (“even numbers,” “increasing by two”).
- Type 2 Processing (Analytic/Reflective): Deliberative, slow, computationally expensive, and governed by cognitive decoupling and counterfactual mental simulation. Type 2 processing is required to question the initial intuitive output of Type 1, recognize the subset problem, and engineer disconfirming tests.
Performance on Shane Frederick’s Cognitive Reflection Test (CRT)—which measures an individual’s disposition to suppress an immediate, compelling, yet incorrect intuitive answer in favor of reflective mathematical verification—is one of the strongest individual-difference predictors of success on the 2-4-6 task. High-CRT individuals possess the metacognitive discipline to pause upon perceiving the salient “increasing by two” pattern, consciously resisting premature closure. Similarly, measures of Need for Cognition (NFC)—a psychometric scale developed by John Cacioppo and Richard Petty that quantifies an individual’s intrinsic motivation to engage in and enjoy effortful cognitive endeavors—correlate positively with sequence variety and testing breadth.
Moreover, the 2-4-6 task serves as a classic demonstration of poor metacognitive calibration. Metacognitive calibration refers to the alignment between an individual’s subjective confidence in their knowledge and their objective factual accuracy. In Wason’s paradigm, individuals who rely exclusively on the Positive Test Strategy exhibit severe overconfidence. Because each confirming trial delivers positive feedback, their subjective probability estimate of being correct approaches 100%. When their rule is rejected, they experience acute psychological disorientation. In contrast, individuals who actively generate disconfirming tests maintain calibrated, realistic assessments of uncertainty throughout their search trajectories.
7.3 Developmental Trajectories and Scientific Literacy
The developmental trajectory of performance on the 2-4-6 task provides profound empirical insights into the ontogeny of human inductive reasoning. Jean Piaget famously posited that during the transition to adolescence (approximately ages 11 to 15), human beings attain the stage of formal operations, characterized by the universal emergence of hypothetico-deductive thought. Developmental studies utilizing child and adolescent adaptations of the 2-4-6 paradigm paint a far more complex picture.
Children in the concrete operational stage (ages 7 to 11) almost universally exhibit unmitigated confirmatory testing. When presented with a numerical seed, they latch onto concrete physical properties and generate uninterrupted streams of positive instances, completely unable to conceptualize why one would ever test a sequence that violates their hypothesis. During adolescence, the capacity to generate alternative hypotheses begins to emerge; however, the spontaneous deployment of genuine disconfirmation remains exceedingly rare without explicit scaffolding.
Surprisingly, educational attainment in adulthood does not automatically eradicate these cognitive vulnerabilities. While undergraduate and postgraduate students enrolled in STEM disciplines (Science, Technology, Engineering, and Mathematics) demonstrate a marginally higher propensity to generate disconfirming tests than humanities students, the baseline rate of spontaneous falsification remains surprisingly low even among professional research scientists. Formal coursework in scientific methodology, statistical inference, and formal propositional logic improves an individual’s theoretical understanding of falsification, but translating that abstract epistemological principle into spontaneous laboratory behavior during an open-ended discovery task requires specialized metacognitive training. Cross-cultural research further indicates that cognitive styles emphasizing holistic thinking versus analytic thinking (as delineated by Richard Nisbett) modulate search patterns: holistic cognitive orientations often lead to faster identification of broad relational rules, whereas analytic orientations frequently lead to protracted perseveration on specific, atomic mathematical rules.
8. Computational and Mathematical Formalizations of the 2-4-6 Task
As cognitive science matured, qualitative debates over whether human beings are “rational” or “irrational” were superseded by rigorous mathematical modeling. Contemporary computational cognitive science has formalized the 2-4-6 task utilizing Bayesian probability theory, rational analysis, and production system architectures, demonstrating that participant behavior can be understood as mathematically optimal inference within a structured environment.
8.1 Bayesian Models of Inductive Inference
The application of Bayesian cognitive modeling—spearheaded by Joshua Tenenbaum, Thomas Griffiths, and Charles Kemp—provides a mathematically elegant framework for understanding how humans learn concepts from sparse data. In a Bayesian formulation, the learner’s goal is to compute the posterior probability distribution $P(H mid D)$ over a vast hypothesis space $\mathcal{H}$, conditioned on observed data $D$ (the sequence of triples and feedback):
$P(H mid D) = \frac{P(D mid H) P(H)}{P(D)} = \frac{P(D mid H) P(H)}{\sum_{H’ in \mathcal{H}} P(D mid H’) P(H’)}$
where $P(H)$ represents the prior probability assigned to a given candidate rule, and $P(D mid H)$ represents the likelihood of observing the data $D$ assuming that hypothesis $H$ is the true governing rule. The true power of the Bayesian model in the 2-4-6 task lies in the formalization of the Size Principle.
Assuming that data points are sampled uniformly at random from the true concept set (the strong sampling assumption), the likelihood of observing a specific triple $x in H$ is inversely proportional to the mathematical size (cardinality or volume) of the hypothesis set $|H|$ raised to the power of the number of observed exemplars $n$:
$P(D mid H) = \left( \frac{1}{|H|} \right)^n$
Now consider the competition between two candidate hypotheses given the initial seed triple 2-4-6: the specific hypothesis $H_{\text{narrow}}$ (“even numbers increasing by two”) and the expansive hypothesis $H_{\text{broad}}$ (“any ascending sequence”). The set of all ascending triples is vast, whereas the set of triples increasing by two is tiny ($|H_{\text{narrow}}| ll |H_{\text{broad}}|$). Consequently, the likelihood of observing the specific exemplar 2-4-6 under $H_{\text{narrow}}$ is immensely higher than observing it under $H_{\text{broad}}$:
$P(\text{2-4-6} mid H_{\text{narrow}}) gg P(\text{2-4-6} mid H_{\text{broad}})$
Under strong sampling, it would be an astronomical coincidence to observe 2, 4, 6 if the true rule were simply “any ascending sequence,” because any numbers like 1-99-1004 or 3-7-82 could have been chosen. Therefore, the Bayesian framework proves that participants who assign an extraordinarily high initial posterior probability to “numbers increasing by two” are behaving with rigorous mathematical rationality. Their initial preference for narrow, highly structured hypotheses is not an irrational cognitive bias; it is the direct, mathematically inevitable consequence of Bayesian updating under the size principle.
8.2 Oaksford and Chater’s Rational Analysis
Building upon the Bayesian framework, Mike Oaksford and Nick Chater developed their pioneering program of Rational Analysis to evaluate human reasoning through the lens of Optimal Data Selection (ODS). Oaksford and Chater argued that human hypothesis testing is designed to maximize expected information gain, quantified formally as the reduction in informational entropy or the Kullback-Leibler (KL) divergence between prior and posterior belief states:
$I(H; D) = \sum_{H in \mathcal{H}} P(H mid D) \log_2 \left( \frac{P(H mid D)}{P(H)} \right)$
Oaksford and Chater demonstrated that when an agent operates in an environment where hypotheses have low prior probabilities and target properties are rare (the rarity assumption), selecting positive test instances ($+tests$) yields a vastly higher Expected Information Gain (EIG) than selecting negative test instances ($-tests$).
When participants in the 2-4-6 task select triples like 8-10-12, they are executing an information-maximization algorithm optimized for real-world sparse environments. The apparent failure of the participants in Wason’s task occurs because Wason constructed a degenerate experimental ecology that radically violates the rarity assumption: the true rule (“ascending numbers”) has an exceptionally high base rate, while the hypothesized rule has a low base rate. Oaksford and Chater’s rational analysis proves that human inductive behavior conforms to optimal informational utility maximization, fully vindicating human bounded rationality against charges of fundamental systemic pathology.
8.3 Algorithmic and Production System Simulations
Beyond mathematical formulations of optimality, cognitive psychologists have constructed mechanistic algorithmic models that simulate the precise cognitive steps executed by participants in the 2-4-6 task. These models are typically implemented within unified cognitive architectures such as ACT-R (Adaptive Control of Thought—Rational), developed by John R. Anderson, or SOAR, developed by Allen Newell and John Laird.
In an ACT-R simulation of the 2-4-6 paradigm, the problem-solving process is modeled as a dynamic interaction between declarative memory chunks (representing numerical knowledge, past instances, and active rules) and procedural production rules (condition-action statements that govern behavior):
- Chunk Retrieval: Upon perceiving 2-4-6, declarative memory retrieves mathematical relation chunks. The chunk for “difference of two” has a high base-level activation due to perceptual priming, causing it to be retrieved first and placed into the goal buffer as the active working hypothesis.
- Production Firing: A production rule matches the current goal buffer: “IF the goal is to test hypothesis H, THEN generate a triple that conforms to H.” This production fires automatically, synthesizing a positive test instance (e.g., 4-6-8).
- Memory Decay and Looping: ACT-R explicitly incorporates the temporal decay of memory traces. Unless a participant explicitly encodes the experimenter’s feedback as a long-term declarative fact, past trials experience activation decay. If a participant’s working memory buffer becomes overloaded, the activation of previously falsified instances fades, causing the production system to retrieve the same or structurally identical hypotheses, accurately simulating the perseverative looping observed in human empirical trials.
These computational production models successfully generate synthetic trial sequences whose quantitative distributions—such as average sequence length, frequency of positive testing, and failure rates—closely mirror the real-world behavioral data gathered from human experimental subjects.
9. Comparative Analysis: 2-4-6 Task versus the Wason Selection Task
Peter Wason is uniquely immortalized in the history of cognitive psychology for inventing not one, but two of the most famous experimental paradigms in the study of human reasoning: the 2-4-6 Hypothesis-Testing Task (1960) and the Four-Card Selection Task (1966). A rigorous comparative analysis of these two experimental monuments reveals profound insights into the bifurcated nature of human rationality across deductive and inductive domains.
9.1 Inductive Generation versus Deductive Selection
The epistemological and structural divergences between the 2-4-6 task and the Selection Task are fundamental:
| Dimension | 2-4-6 Hypothesis-Testing Task (1960) | Four-Card Selection Task (1966) |
|---|---|---|
| Epistemological Domain | Inductive Reasoning: Open-ended discovery of an unknown universal law from sparse empirical exemplars. | Deductive Reasoning: Evaluation of whether an explicit conditional rule ($P \rightarrow Q$) is violated. |
| Problem Space | Infinite / Open-Ended: The participant can generate any combination of numbers within $\mathbb{R}^3$. | Finite / Constrained: The participant is presented with exactly four physical or symbolic cards ($P, neg P, Q, neg Q$). |
| Cognitive Action | Generative: Active construction of novel empirical data points and hypotheses. | Evaluative / Selective: Passive inspection and choice of pre-existing stimuli. |
| Underlying Cognitive Driver | Positive Test Strategy (PTS): Testing instances expected to possess the target property. | Matching Bias & Verification: Selecting cards mentioned in the conditional statement ($P$ and $Q$). |
| Nature of Failure | Failure to transcend an embedded subset relationship via complementary testing. | Failure to apply the formal logical rule of modus tollens (selecting the $neg Q$ card). |
In the Selection Task, participants are presented with a conditional rule—for example: “If a card has a vowel on one side, then it has an even number on the other side” ($P \rightarrow Q$)—alongside four cards showing, for example, A ($P$), B ($neg P$), 4 ($Q$), and 7 ($neg Q$). The logically correct selection required to establish validity is to turn over the A card (modus ponens) and the 7 card (modus tollens). However, typical participants overwhelmingly select A and 4 ($P$ and $Q$), attempting to confirm the rule by observing a vowel behind the even number, while completely neglecting the falsifying potential of the $neg Q$ card (7).
While the Selection Task tests whether individuals understand the formal logical necessity of falsification when evaluating an existing rule, the 2-4-6 task tests how individuals behave when they must generate both the rules and the tests themselves. Thus, the 2-4-6 task engages far more complex metacognitive, generative, and search mechanisms than the Selection Task.
9.2 The Role of Context and Deontic Content
One of the most spectacular discoveries in the history of the Selection Task was the thematic content effect. When the abstract letters and numbers of the Selection Task are replaced with real-world social or deontic rules—such as “If a person is drinking beer, then that person must be over 21 years of age”—participant performance undergoes a breathtaking transformation: correct falsifying selections ($P$ and $neg Q$) skyrocket from less than 10% to over 80%. Evolutionary psychologists such as Leda Cosmides and John Tooby argued that this dramatic facilitation reveals the existence of an evolved, domain-specific neurocognitive module: a specialized “cheater detection mechanism” designed to identify individuals who take a benefit without fulfilling a mandatory social cost. Alternatively, Patricia Cheng and Keith Holyoak explained this facilitation through Pragmatic Reasoning Schemas—generalized cognitive frameworks developed through everyday social experience governing permissions and obligations.
Crucially, this dramatic content facilitation effect fails to transfer cleanly to the 2-4-6 task. Because the 2-4-6 task is inherently a problem of inductive mathematical categorization, researchers have found it exceptionally difficult to construct “deontic” equivalents that naturally trigger cheater-detection or permission schemas. Transforming the 2-4-6 task into thematic narratives (e.g., discovering the rules of an alien ecosystem or uncovering the culinary preferences of a monarch) often fails to eliminate the Positive Test Strategy, because human beings naturally explore novel semantic spaces by looking for positive exemplars of categories. This stark contrast highlights the unique epistemological status of the 2-4-6 task: it interrogates pure inductive hypothesis generation, a domain largely insulated from social-contract reasoning modules.
9.3 Common Cognitive Underpinnings
Despite their epistemological differences, both of Wason’s tasks expose deep, shared cognitive mechanisms that define modern Dual-Process Theory:
- Matching Bias and Perceptual Salience: In both tasks, participants are cognitively seduced by superficial perceptual features explicitly stated in the problem prompt. In the Selection Task, participants pick the cards named in the conditional rule (matching bias). In the 2-4-6 task, participants anchor exclusively onto the salient arithmetic relationships present in the seed exemplar 2-4-6.
- Neglect of Negative Space: Both paradigms expose a profound cognitive blindness to the “negative space” of logical problems. Human thinkers struggle fundamentally to conceptualize non-events, non-instances, and counterfactual possibilities. In the card task, they ignore the $neg Q$ card; in the number task, they refuse to generate sequences that fall into $\bar{H}$.
- System 1 Dominance over System 2: In both paradigms, autonomous, heuristic-driven System 1 processes deliver immediate, highly intuitive solutions that feel compellingly correct. Unless an individual possesses the metacognitive capacity and cognitive motivation to engage computationally expensive System 2 overrides, the intuitive solution is immediately endorsed, producing systemic cognitive error across both deductive and inductive domains.
10. Philosophical and Epistemological Ramifications
The empirical findings of the 2-4-6 task transcend experimental psychology, striking directly at the foundational core of Western philosophy of science and epistemology. By operationalizing the interaction between human cognitive heuristics and abstract hypothesis spaces, the paradigm illuminates centuries of philosophical debate concerning the nature of scientific progress, the problem of induction, and the sociological persistence of ideological dogmatism.
10.1 The Psychology of Scientific Discovery
Karl Popper’s normative philosophy asserted that science progresses exclusively through bold conjecture and relentless falsification. Yet, when historians and sociologists of science inspect the actual history of scientific discovery, they encounter an empirical reality that mirrors the 2-4-6 paradigm far more than Popper’s idealized standard.
In his monumental 1962 work, The Structure of Scientific Revolutions, Thomas Kuhn described the protracted periods of “normal science” that define scientific disciplines. During normal science, researchers do not attempt to falsify the reigning theoretical paradigm. Instead, they engage in focused “puzzle-solving”—a process that is functionally identical to the Positive Test Strategy. Working within the established paradigm, scientists formulate specific, nested hypotheses that take the paradigm’s core assumptions for granted, meticulously generating experiments designed to confirm, articulate, and refine those theories. When an experiment yields negative or anomalous data, scientists rarely discard their primary theory; rather, they treat the result as an experimental error, an instrument malfunction, or an anomaly to be resolved later.
Philosopher of science Imre Lakatos extended this critique through his Methodology of Scientific Research Programmes. Lakatos demonstrated that successful scientific programmes possess a “hard core” of fundamental theoretical assumptions that are deliberately shielded from refutation by a vast “protective belt” of auxiliary hypotheses. When anomalies strike, it is the auxiliary hypotheses that are modified, altered, or abandoned, preserving the hard core intact.
The 2-4-6 task provides the psychological micro-foundations for Kuhn’s and Lakatos’s observations. The reluctance of participants to abandon their narrow arithmetic hypotheses in the face of ambiguity reflects the deeply ingrained human tendency to protect theoretical hard cores. Furthermore, cognitive psychologists like Kevin Dunbar, utilizing in vivo laboratory studies of real-world molecular biology laboratories, have revealed that practicing scientists overwhelmingly deploy the Positive Test Strategy. Scientists intentionally run tests expected to work, using positive testing to build stable experimental baselines. Genuine falsification is deployed primarily as an exceptional, high-stakes strategic pivot when a paradigm is completely exhausted. Thus, Wason’s task exposes the psychological impossibility of pure Popperianism as an everyday operational methodology.
10.2 The Problem of Induction Revisited
The 2-4-6 paradigm offers an empirical realization of the profound epistemological dilemmas first articulated by David Hume and Nelson Goodman.
In the eighteenth century, David Hume articulated the classic Problem of Induction: no amount of empirical observation can ever provide rational justification for the uniformity of nature or the truth of universal claims. We assume that the sun will rise tomorrow because it has risen every day in recorded history, but this inference rests on the circular assumption that the future will necessarily conform to the past. In the 2-4-6 task, participants experience Hume’s dilemma in microcosm: they observe that 4-6-8, 10-12-14, and 100-102-104 all return positive feedback, and they infer with absolute subjective certainty that the universal rule is “increasing by two.” They conflate empirical consistency with rational proof, falling victim to the classic inductivist illusion.
In 1955, philosopher Nelson Goodman introduced the New Riddle of Induction, demonstrating that the challenge is not merely justifying induction, but determining which predicates are *projectable*. Goodman introduced the arbitrary predicate “grue” (an object is grue if it is observed to be green before time $T$, and blue thereafter). All empirical observations of green emeralds before time $T$ provide equal confirmatory evidence for the hypothesis that “all emeralds are green” and the hypothesis that “all emeralds are grue.” How do we rationally choose between these competing generalizations?
The 2-4-6 task is the exact psychological instantiation of Goodman’s riddle. Given the seed 2-4-6, the data provides equal empirical corroboration for an infinite variety of nested predicates: “increasing by two,” “consecutive even integers,” “numbers less than 1000,” and “ascending numbers.” The human mind resolves Goodman’s paradox through entrenchment: we instinctively project predicates that are familiar, salient, and historically reinforced in our cognitive development (such as standard arithmetic progressions). However, the 2-4-6 task exposes the grave danger of cognitive entrenchment: when reality is governed by an unentrenched, maximally broad predicate (“any ascending sequence”), our cognitive entrenchments blind us to the true structure of the world.
10.3 Confirmation Bias in Societal and Political Reasoning
While the 2-4-6 task operates within an abstract domain of numbers, the cognitive dynamics it uncovers—specifically the Positive Test Strategy operating within embedded subsets—provide a terrifyingly accurate explanatory model for the sociological phenomena of ideological polarization, echo chambers, and the proliferation of conspiratorial beliefs.
In modern sociopolitical discourse, an individual’s ideological framework functions as a hypothesized rule $H$. When an ideologically committed individual seeks information, they almost exclusively deploy the Positive Test Strategy: they seek out news sources, social media algorithms, and peer networks that they expect will confirm their worldview ($x in H$). In a complex, noisy societal environment containing trillions of informational data points, the individual will effortlessly discover millions of positive exemplars that corroborate their narrative (e.g., examples of political corruption, economic malfeasance, or cultural decline). Each positive exemplar returns categorical “conforms” feedback from their social bubble, driving subjective confidence toward absolute certainty.
The epistemic pathology of the social media echo chamber is mathematically isomorphic to the 2-4-6 task. The ideologue never executes a negative test strategy: they never deliberately immerse themselves in competing intellectual ecosystems, nor do they seek out empirical data specifically designed to test the boundary conditions or potential falsity of their deepest beliefs. Furthermore, when disconfirming evidence is unavoidably encountered, the cognitive dissonance it produces is so psychologically distressing that the individual immediately engages in pseudo-falsification or constructs ad-hoc auxiliary hypotheses to preserve their ideological core. The 2-4-6 task demonstrates that dogmatism is not necessarily born of malice, low intelligence, or emotional instability; it is the natural, predictable outcome of deploying standard human cognitive heuristics within structurally biased informational environments.
11. Neurocognitive and Clinical Perspectives
The transition of cognitive psychology into cognitive neuroscience and neuropsychiatry has enabled researchers to map the neuroanatomical substrates, electrophysiological correlates, and clinical manifestations associated with hypothesis generation, feedback processing, and cognitive flexibility in the 2-4-6 task.
11.1 Neuroimaging Studies of Hypothesis Generation and Disconfirmation
Functional Magnetic Resonance Imaging (fMRI) studies investigating inductive reasoning tasks modeled on the 2-4-6 paradigm reveal a distributed, highly coordinated frontoparietal cognitive control network responsible for navigating hypothesis spaces:
- Dorsolateral Prefrontal Cortex (dlPFC): The dlPFC, particularly in the left hemisphere, exhibits intense blood-oxygen-level-dependent (BOLD) activation during the active generation of hypotheses and the strategic planning of test sequences. The dlPFC is fundamentally responsible for maintaining abstract relational rules in working memory and orchestrating cognitive decoupling.
- Ventrolateral Prefrontal Cortex (vlPFC): The vlPFC plays an indispensable role in rule selection and the active inhibition of prepotent perceptual responses. To generate a disconfirming negative test like 1-2-10, the vlPFC must vigorously suppress the salient, intuitive impulse to test the arithmetic pattern “+2.”
- Dorsal Anterior Cingulate Cortex (dACC): The dACC functions as the brain’s premier conflict-detection and error-monitoring hub. When an individual who firmly believes their hypothesis is correct receives the unexpected feedback “does not conform,” or conversely receives an unexpected “conforms” to a test designed to fail, the dACC exhibits a massive surge in activation, signaling a profound prediction error and signaling the dlPFC to initiate behavioral adaptation.
- The Striatum and Ventral Tegmental Area (VTA): The mesolimbic dopamine system plays a decisive, often subversive role in hypothesis testing. Confirmatory feedback (“conforms”) triggers robust phasic dopamine release in the nucleus accumbens and striatum, providing neurochemical reinforcement. This reward signal explains the seductive psychological comfort of the Positive Test Strategy: human beings experience positive validation as an intrinsically rewarding neurochemical event, reinforcing the drive to seek confirming data.
Electrophysiological studies utilizing Event-Related Potentials (ERPs) reveal precise temporal dynamics during 2-4-6 performance. The reception of disconfirming feedback elicits a sharp, negative voltage deflection approximately 250 milliseconds post-stimulus: the Feedback-Related Negativity (FRN), originating in the anterior cingulate. This is rapidly followed by the P300 (specifically the P3b) component, a large positive deflection reflecting the cognitive updating of mental models in working memory. Individuals who successfully solve the task exhibit significantly larger P3b amplitudes following negative feedback than individuals who perseverate, indicating that successful solvers immediately allocate heavy attentional resources to cognitive restructuring.
11.2 Cognitive Biases in Medical Diagnosis and Clinical Judgments
The cognitive dynamics of the 2-4-6 task carry profound, life-or-death implications in the domain of clinical medicine. Medical error researchers have identified diagnostic premature closure—the termination of the diagnostic search process before the correct diagnosis has been definitively established—as the single most common cognitive error leading to catastrophic clinical misdiagnoses.
When a physician evaluates a patient presenting with complex, ambiguous symptoms, the physician’s initial clinical impression represents a candidate hypothesis $H$. Under the influence of the Positive Test Strategy, the clinician frequently orders laboratory tests, imaging studies, and specialized consultations specifically designed to confirm their suspected diagnosis ($+tests$). For example, if a physician suspects pneumonia, they order tests looking for pneumonia. If those tests return positive, the physician prematurely declares the diagnosis, often failing to recognize that the patient’s symptoms are actually driven by an underlying, far broader pathology—such as systemic lupus erythematosus or a rare pulmonary malignancy—that completely encompasses the symptoms of pneumonia (an embedded subset relationship).
To combat this deadly clinical manifestation of the 2-4-6 trap, modern medical education has increasingly integrated diagnostic debiasing frameworks. Medical trainees are explicitly trained in differential diagnosis elimination—a formal clinical instantiation of the negative test strategy. Physicians are instructed to actively ask: “What is the worst possible alternative diagnosis that could explain these symptoms, and what specific clinical test can I order that has the diagnostic power to rule it out?”
11.3 Neuropsychological and Psychopathological Variations
Clinical neuropsychology provides critical insights into the modular architecture of hypothesis testing through the examination of clinical populations suffering from specific neurocognitive impairments:
- Frontal Lobe Lesion Patients: Individuals who have suffered focal lesions to the prefrontal cortex—particularly the ventromedial or dorsolateral sectors—exhibit catastrophic perseveration on the 2-4-6 task. When administered the task, they frequently latch onto the very first salient mathematical pattern (e.g., “+2”) and generate endless streams of positive tests. Even when the experimenter explicitly rejects their rule, they will state: “I’ll try again,” and immediately proceed to test 20-22-24, completely incapable of shifting cognitive set. Their behavior closely mirrors their performance on the Wisconsin Card Sorting Test (WCST), demonstrating that intact prefrontal machinery is mandatory for cognitive flexibility.
- Schizophrenia and the Jumping to Conclusions (JTC) Bias: Patients with schizophrenia, particularly those experiencing active persecutory delusions, exhibit a well-documented cognitive deficit known as the Jumping to Conclusions bias, typically measured via the Beads Task. In the 2-4-6 task, individuals with schizophrenia generate significantly fewer triples than neurotypical controls prior to declaring their definitive rule. They observe 2-4-6, test perhaps a single confirming triple (e.g., 8-10-12), and immediately declare the rule with absolute subjective certainty. This hyper-rapid hypothesis acceptance is driven by aberrant salience—the inappropriate assignment of intense significance to neutral or routine stimuli.
- Obsessive-Compulsive Disorder (OCD): Patients diagnosed with OCD display the exact opposite behavioral pathology. Driven by an overwhelming intolerance of uncertainty and a profound deficit in subjective cognitive closure, individuals with OCD will generate dozens, sometimes hundreds of triples without ever announcing a rule. Even when they have fully mapped the boundaries of the latent rule through extensive positive and negative testing, they remain crippled by epistemic doubt, unable to reach the internal threshold of subjective certainty required to terminate the task.
- Autism Spectrum Disorder (ASD): Individuals with ASD often exhibit unique performance profiles on the 2-4-6 task, driven by heightened systemizing tendencies. Because individuals with ASD frequently possess exceptional mathematical pattern-recognition skills, they often formulate hyper-complex algebraic hypotheses. However, their lower susceptibility to social-pragmatic framing effects (Gricean implicatures) can occasionally provide an unexpected advantage: they are less seduced by the communicative intent of the experimenter, allowing some autistic individuals to identify broad relational rules with remarkable efficiency once arithmetic avenues are systematically eliminated.
12. Pedagogical Implementations and Debiasing Frameworks
Given its unparalleled elegance and demonstrable capacity to expose systemic cognitive vulnerabilities, the Wason 2-4-6 task has been extensively adopted as a foundational pedagogical tool across secondary, undergraduate, and executive education. Furthermore, the paradigm has served as the empirical testbed for engineering institutional decision architectures designed to inoculate professional intelligence analysts, corporate leaders, and scientific researchers against catastrophic confirmation traps.
12.1 Teaching Scientific Methodology through the 2-4-6 Task
In academic pedagogical settings, the 2-4-6 task is widely utilized as an immersive, experiential classroom demonstration. Rather than passively lecturing students on the abstract epistemological principles of Karl Popper, Francis Bacon, or David Hume, instructors administer the 2-4-6 task as an interactive, live experiment during the opening minutes of a course on research methodology, scientific inquiry, or cognitive psychology.
The experiential impact of the demonstration follows a standardized pedagogical arc:
- The instructor writes “2, 4, 6” on the board and invites students to discover the rule by proposing triples and recording the feedback.
- The classroom inevitably devolves into enthusiastic confirmation: students eagerly shout out “8, 10, 12,” “20, 22, 24,” and “100, 102, 104,” with the instructor dutifully recording “Yes” on the board.
- When a student confidently declares: “The rule is numbers increasing by two,” and the instructor announces that this rule is incorrect, a collective shock sweeps through the room.
- The instructor uses this acute cognitive dissonance as a profound teaching moment, demonstrating to students that their entire educational history has conditioned them to seek validation rather than refutation.
By forcing students to log their internal reasoning processes across subsequent iterations, instructors guide them toward the profound realization that an empirical hypothesis can only be validated if the researcher actively attempts to break it. Longitudinal empirical studies tracking educational outcomes indicate that students who undergo experiential 2-4-6 training exhibit a significantly higher propensity to construct rigorous experimental control groups and deploy disconfirming controls in subsequent scientific laboratory coursework.
12.2 Cognitive Debiasing Interventions and Decision Architecture
In high-stakes corporate, military, and intelligence environments, cognitive biases cannot be cured through simple self-awareness. Recognizing this limitation, organizational theorists and cognitive psychologists have developed structured analytic techniques designed to hardwire the mechanics of the Negative Test Strategy into institutional decision architectures.
The most prominent of these frameworks is the Analysis of Competing Hypotheses (ACH), developed by CIA veteran Richards Heuer. Heuer designed ACH specifically to overcome the intelligence community’s pervasive vulnerability to confirmation bias and premature diagnostic closure. ACH forces analysts to construct an explicit matrix displaying all conceivable hypotheses across the top and all available intelligence data down the side. Crucially, the analyst is prohibited from evaluating how well the data confirms a favored hypothesis. Instead, the mathematical algorithm of ACH forces the analyst to evaluate the data strictly in terms of inconsistency, methodically working to eliminate hypotheses that are contradicted by the evidence. ACH is, in essence, an institutionalized operationalization of Tweney’s DAX/MED and Wason’s negative testing strategies.
Similarly, high-performing strategic organizations employ institutionalized adversarial testing through Red Teaming. A designated Red Team is formally tasked with assuming the truth of the opposite: their explicit mandate is to find data, scenarios, and operational conditions under which the organization’s primary strategic plan will experience catastrophic failure. In modern software engineering and intelligence forecasting platforms, automated algorithmic prompts now force users to specify their falsification criteria in advance: before a strategic forecast is locked in, the system requires the analyst to answer: “What specific future empirical event would prove that your current assessment is definitively wrong?”
12.3 Future Directions in Inductive Reasoning Research
As cognitive science enters the era of artificial intelligence and complex computational systems, the Wason 2-4-6 paradigm continues to inspire cutting-edge empirical research. One of the most active contemporary frontiers involves the evaluation of Large Language Models (LLMs)—such as OpenAI’s GPT-4, Anthropic’s Claude, and Google’s Gemini—on inductive reasoning benchmarks modeled on the 2-4-6 task.
Recent experiments testing advanced artificial intelligence models on the 2-4-6 paradigm have revealed a fascinating emergent phenomenon: when prompted to discover the latent rule, LLMs exhibit cognitive behaviors remarkably isomorphic to human participants. Because LLMs are trained on massive corpora of human text that heavily reflect natural human conversational implicatures and the Positive Test Strategy, they frequently fall into the exact same embedded subset traps as human undergraduates, generating strings of positive, confirming triples (e.g., 8-10-12) and prematurely announcing “numbers increasing by two.” However, when prompted with specialized metacognitive instructions (such as Chain-of-Thought prompting or explicit commands to execute competitive hypothesis elimination), their success rates dramatically diverge, providing a rich computational testbed for modeling the boundaries of artificial inductive intelligence.
Furthermore, contemporary researchers are extending the 2-4-6 paradigm into probabilistic, non-linear, and dynamic environments. In these advanced variations, experimenter feedback is no longer deterministic; a conforming triple might receive a “conforms” response only 80% of the time (stochastic feedback), or the governing mathematical rule might slowly drift over time (non-stationary environments). Investigating how human agents and multi-agent computational systems navigate such noisy, complex hypothesis spaces ensures that Peter Wason’s elegant, seventy-year-old numerical puzzle remains at the absolute cutting edge of cognitive psychology, neuroscience, and artificial intelligence.
In the final analysis, Peter Cathcart Wason’s 2-4-6 hypothesis-testing task stands as an enduring monument to the power of experimental minimalism. With nothing more than a pencil, a sheet of paper, and three ordinary numbers, Wason peeled back the superficial veneer of human intellectual arrogance, exposing the profound, systemic tensions that exist between human cognitive heuristics and the uncompromising demands of objective truth. Whether viewed as an indictment of human irrationality, a testament to the ecological brilliance of the Positive Test Strategy, or a mathematical warning against the seductive comfort of confirmatory echo chambers, the lesson of 2-4-6 remains eternal: true understanding is never attained by celebrating the evidence that confirms our beliefs; it is won through the intellectual courage to seek the evidence that shatters them.
References
Bruner, J. S., Goodnow, J. J., & Austin, G. A. (1956). A study of thinking. John Wiley & Sons.
Cheng, P. W., & Holyoak, K. J. (1985). Pragmatic reasoning schemas. Cognitive Psychology, 17(4), 391–416. https://doi.org/10.1016/0010-0285(85)90014-3
Cosmides, L. (1989). The logic of social exchange: Has natural selection shaped how humans reason? Studies with the Wason selection task. Cognition, 31(3), 187–276. https://doi.org/10.1016/0010-0277(89)90023-1
Ericsson, K. A., & Simon, H. A. (1993). Protocol analysis: Verbal reports as data (Rev. ed.). MIT Press.
Evans, J. S. B. (1989). Bias in human reasoning: Causes and consequences. Lawrence Erlbaum Associates.
Evans, J. S. B. (2006). The heuristic-analytic theory of reasoning: Extension and evaluation. Psychonomic Bulletin & Review, 13(3), 378–395. https://doi.org/10.3758/BF03193858
Frederick, S. (2005). Cognitive reflection and decision making. Journal of Economic Perspectives, 19(4), 25–42. https://doi.org/10.1257/089533005775196732
Goodman, N. (1955). Fact, fiction, and forecast. Harvard University Press.
Grice, H. P. (1975). Logic and conversation. In P. Cole & J. L. Morgan (Eds.), Syntax and semantics: Vol. 3. Speech acts (pp. 41–58). Academic Press.
Heuer, R. J. (1999). Psychology of intelligence analysis. Center for the Study of Intelligence.
Kahneman, D. (2011). Thinking, fast and slow. Farrar, Straus and Giroux.
Klayman, J., & Ha, Y.-W. (1987). Confirmation bias re-examined. Psychological Review, 94(2), 211–228. https://doi.org/10.1037/0033-295X.94.2.211
Klayman, J., & Ha, Y.-W. (1989). Hypothesis testing in rule discovery: Strategy, structure, and content. Journal of Experimental Psychology: Learning, Memory, and Cognition, 15(4), 596–604. https://doi.org/10.1037/0278-7393.15.4.596
Kuhn, T. S. (1962). The structure of scientific revolutions. University of Chicago Press.
Lakatos, I. (1978). The methodology of scientific research programmes: Philosophical papers (Vol. 1) (J. Worrall & G. Currie, Eds.). Cambridge University Press.
Lord, C. G., Lepper, M. R., & Preston, E. (1984). Considering the opposite: A corrective strategy for social judgment. Journal of Personality and Social Psychology, 47(6), 1231–1243. https://doi.org/10.1037/0022-3514.47.6.1231
Oaksford, M., & Chater, N. (1994). A rational analysis of the selection task as optimal data selection. Psychological Review, 101(4), 608–631. https://doi.org/10.1037/0033-295X.101.4.608
Oaksford, M., & Chater, N. (2003). Optimal data selection and the probabilistic approach to human reasoning. In D. Hardman & L. Maflitt (Eds.), Thinking: Psychological perspectives on reasoning, judgment and decision making (pp. 51–72). John Wiley & Sons.
Popper, K. R. (1959). The logic of scientific discovery. Hutchinson & Co.
Stanovich, K. E. (2011). Rationality and the reflective mind. Oxford University Press.
Tenenbaum, J. B., & Griffiths, T. L. (2001). Generalization, similarity, and Bayesian inference. Behavioral and Brain Sciences, 24(4), 629–640. https://doi.org/10.1017/S0140525X01000061
Tweney, R. D., Doherty, M. E., Worner, W. J., Pliske, D. B., Mynatt, C. R., Gross, K. A., & Arkkelin, D. L. (1980). Strategies of rule discovery on an inductive task. Quarterly Journal of Experimental Psychology, 32(1), 109–123. https://doi.org/10.1080/14640748008401146
Wason, P. C. (1960). On the failure to eliminate hypotheses in a conceptual task. Quarterly Journal of Experimental Psychology, 12(3), 129–140. https://doi.org/10.1080/17470216008416717
Wason, P. C. (1966). Reasoning. In B. M. Foss (Ed.), New horizons in psychology (pp. 135–151). Penguin Books.
Wason, P. C. (1968). Reasoning about a rule. Quarterly Journal of Experimental Psychology, 20(3), 273–281. https://doi.org/10.1080/14640746808400161