In the landscape of modern cognitive neuroscience, the transition from descriptive psychological taxonomy to formal mathematical architecture represents one of the most consequential paradigm shifts of the twenty-first century. At the epicenter of this transformation stands P. Read Montague, whose foundational contributions to computational neurobiology have radically reformulated our comprehension of decision-making under uncertainty. Historically, normative economic theories treated the human decision-maker as a static rational agent operating within stationary probability distributions—an idealized construct that repeatedly buckled when confronted with empirical human behavior. Montague and his contemporaries recognized that biological brains did not evolve to solve closed-form optimization problems in static environments; rather, they evolved to harvest resources, evade threats, and negotiate survival within deeply uncertain, dynamically fluctuating, and non-stationary worlds.
Central to this ecological imperative is the explore-exploit trade-off, a fundamental computational dilemma wherein an agent must constantly adjudicate between monetizing known reward structures and sacrificing immediate utility to sample uncertain alternatives that may yield superior long-term returns. To dissect the precise neurocomputational mechanisms orchestrating this dynamic arbitration, Montague devised experimental paradigms that broke sharply from standard laboratory multi-armed bandits. Among these, the Leapfrog Task stands as a singularly brilliant experimental vehicle. Unlike canonical stationary or Gaussian random-walk paradigms, the Leapfrog Task introduces structured non-linear value dynamics where dormant, unchosen options intermittently “leapfrog” over active options according to hidden, non-stationary transition rules. Navigating this landscape requires an agent to maintain complex counterfactual state estimates, model hidden environmental volatility, and deploy strategically timed exploratory excursions based on epistemic curiosity rather than simple stochastic motor variability.
However, delineating the neural correlates of these computational operations using functional magnetic resonance imaging (fMRI) presents an extraordinary methodological challenge. Task-based fMRI operates at the uncomfortable intersection of high-dimensional, millisecond-scale cognitive computations and the sluggish, non-linear biophysics of the blood-oxygen-level-dependent (BOLD) hemodynamic response function. Furthermore, the essential prefrontal, cingulate, and subcortical nodes that theoretically implement dynamic value updating—most notably the ventromedial prefrontal cortex, the frontopolar cortex, and midbrain dopaminergic nuclei—are precisely the territories plagued by macroscopic magnetic susceptibility artifacts, spatial distortions, and physiological noise. This comprehensive analysis unpacks the theoretical, algorithmic, and neuroimaging dimensions of Montague’s Leapfrog Task, charting how model-based fMRI bridges latent cognitive parameters with human neuroanatomy, illuminates the circuitry of exploratory choice, and redefines our understanding of psychiatric vulnerability.
1. Theoretical Foundations of Read Montague’s Computational Neurobiology
1.1 The Evolution of Neuroeconomic Frameworks in Decision Science
The dawn of neuroeconomics was catalyzed by the realization that classical normative economics failed to account for the physical constraints and evolutionary history of biological wetware. Early neoclassical paradigms, typified by expected utility theory and axiomatic revealed preferences, operated on the assumption of infinite computational capacity and static utility surfaces. In stark contrast, Read Montague and colleagues at the interface of theoretical physics, computer science, and molecular neurobiology advanced an epistemological model predicated on biologically grounded algorithmic principles. This shift reconceptualized the biological brain not as a passive utility calculator, but as an active, thermodynamic inference engine designed to minimize surprise, conserve energetic resources, and extract regularities from noisy ecological niches.
Drawing heavily from Horace Barlow’s efficient coding hypothesis and the emerging tenets of predictive processing, Montague posited that sensory and reward-processing systems are fundamentally tuned to maximize information transmission under metabolic constraints. In reward-guided navigation, this principle mandates that neural circuits do not encode absolute objective value; instead, they encode deviations from internally generated predictions. The integration of temporal difference reinforcement learning (TDRL) algorithms into neurobiology—spearheaded by Montague, Peter Dayan, and Wolfram Schultz—demonstrated that midbrain dopaminergic activity tracks reward prediction errors (RPEs), fundamentally bridging cellular biophysics with computational reinforcement learning. By framing dynamic behavioral testing through this algorithmic lens, Montague established the foundation of computational phenotyping, a diagnostic and investigative approach that maps high-level behavioral divergence directly onto mathematically parameterized biophysical circuits.
Computational phenotyping recognizes that overt behavioral choices are merely the low-dimensional readouts of complex, latent computational trajectories. In the context of decision science, two individuals displaying identical macroscopic choice frequencies on a behavioral assay may arrive at those choices through radically distinct computational routes: one driven by accelerated learning rates operating over short temporal horizons, the other governed by inflated baseline expectations coupled with hyper-sensitive risk aversion. Montague’s algorithmic frameworks formalize these invisible trajectories, establishing continuous state spaces where latent parameters—such as belief precision, counterfactual drift rates, and information exploration bonuses—can be extracted and mapped onto underlying neural substrates via functional neuroimaging.
1.2 Conceptualization of the Leapfrog Task Paradigm
To rigorously interrogate exploratory decision-making, Montague engineered the Leapfrog Task as an explicit departure from traditional multi-armed bandit paradigms. In canonical four-armed bandit tasks, the payout values of distinct choices typically evolve according to independent, continuous Gaussian random walks. While such environments successfully induce uncertainty, they allow agents to rely on passive, model-free associative learning or simple temporal decay heuristics. An agent tracking a slowly drifting Brownian motion need only monitor the recent history of local rewards to sustain a moderately successful policy, minimizing the necessity for complex internal models of environmental state transitions.
The Leapfrog Task introduces an entirely different class of non-stationary, multi-alternative decision mechanics. In this architecture, the generative model specifies that while an agent exploits a known, high-value alternative, the latent, unchosen options do not merely drift passively; instead, their underlying reward distributions accumulate latent potential or undergo discrete, non-linear state jumps. At mathematically governed intervals, an unobserved option “leapfrogs” the currently exploited option, instantaneously claiming superior utility. Crucially, the occurrence of this leapfrog event cannot be directly observed without actively sacrificing the assured payout of the incumbent option to visually or numerically sample the competing alternatives.
This design creates a profoundly non-linear value topology. If an agent remains anchored to an exploitative policy for too long, the relative opportunity cost accelerates exponentially as the hidden state of an unchosen target undergoes a leapfrog shift. Conversely, premature or hyper-frequent sampling of alternative options degrades immediate reward harvest due to the transaction costs and lower baseline payouts of dormant options. Montague devised this architecture to capture the true ecological essence of foraging and competitive predation, wherein resources deplete locally while regenerating distally, compelling biological organisms to build hierarchical predictive representations of the environment rather than relying on reactive reinforcement.
1.3 Differentiating Directed Exploration from Random Exploration
A foundational theoretical imperative within computational neurobiology is the formal segregation of directed exploration from random exploration. In classic reinforcement learning architectures, exploration is often implemented heuristically as a stochastic perturbation of choice policies. The most ubiquitous instantiation is the softmax (or Boltzmann) action-selection rule, wherein an agent computes the expected value of available actions and selects among them probabilistically as a function of an inverse temperature parameter, denoted as $\beta$. When $\beta$ approaches zero, choice selection degrades into pure random exploration, driven by endogenous thermal or neural noise. While mathematically convenient, this mechanism is computationally agnostic: the agent allocates exploratory capital uniformly across options without regard to the magnitude of uncertainty surrounding any specific alternative.
In contrast, directed exploration represents a deliberate, teleological allocation of cognitive resources toward sources of epistemic ambiguity. Under directed exploration frameworks, agents compute an explicit “information bonus” that inflates the subjective utility of poorly characterized states. This can be formalized through algorithms such as Upper Confidence Bound (UCB) strategies or via Bayesian active inference, wherein choice selection is guided by the minimization of expected free energy. In an active inference formulation, actions are selected not merely to harvest pragmatic value (reward), but to maximize epistemic value (information gain), systematically resolving uncertainty within the agent’s internal generative model of the world.
Within Montague’s Leapfrog Task, this distinction becomes analytically tractable. Because unchosen options evolve latently, the Bayesian variance (uncertainty) surrounding their true value distribution expands monotonically with every consecutive trial they remain unsampled. An agent executing directed exploration systematically samples an unchosen option precisely when the integrated uncertainty surrounding its potential leapfrog breach crosses a critical epistemic threshold. By deploying high-density computational behavioral models, investigators can mathematically disentangle the behavioral signatures of true directed exploration—quantified via Bayesian belief distributions and information bonuses—from stochastic motor noise or lapse rates, revealing the dedicated prefrontal circuits that arbitrate epistemic curiosity.
2. Mathematical Formulation and Task Architecture of the Leapfrog Paradigm
2.1 State Space Dynamics and Value Transition Matrices
The operational engine of the Leapfrog Task is formally characterized as a non-stationary Markov Decision Process (MDP) augmented by partially observable environmental states. Let the environment consist of a finite set of discrete choice options $\mathcal{A} = {a_1, a_2, dots, a_K}$, where at any discrete trial $t in {1, 2, dots, T}$, each action yields an underlying continuous scalar reward $r_t(a) in \mathbb{R}$. Unlike stationary MDPs where the transition probability matrix $\mathcal{P}(s’ mid s, a)$ is time-invariant, the Leapfrog environment embeds dynamic transition probabilities governing discrete, latent state transitions that continuously recalibrate the expected reward vectors.
The generative model dictates that the currently exploited option $a_{\text{active}}$ provides a steady or slowly depleting mean reward, modeled as:
$$\mu_t(a_{\text{active}}) = \mu_{t-1}(a_{\text{active}}) – \delta_{\text{decay}} + \epsilon_{\text{exploit}}$$
where $\delta_{\text{decay}} ge 0$ represents a predictable linear or exponential depletion rate, and $\epsilon_{\text{exploit}} \sim \mathcal{N}(0, \sigma_{\text{exploit}}^2)$ introduces Gaussian environmental noise. Concurrently, the latent expected rewards of the unchosen, dormant alternatives $a_{\text{dormant}} in \mathcal{A} setminus {a_{\text{active}}}$ progress according to a piecewise stochastic trajectory. Across consecutive unselected trials, these options experience an incremental baseline drift punctuated by a discrete leapfrog jump probability, denoted as $P_{\text{leap}}$. The conditional state transition governing an unchosen alternative’s latent mean can be formalized as:
$$\mu_t(a_k) = \begin{\cases} \mu_{t-1}(a_k) + \Delta_{\text{ju\mp}} + \epsilon_{\text{ju\mp}}, & \text{with probability } P_{\text{leap}}(t – \tau_k) \ \mu_{t-1}(a_k) + \epsilon_{\text{drift}}, & \text{with probability } 1 – P_{\text{leap}}(t – \tau_k) \end{\cases}$$
Here, $\tau_k$ denotes the exact trial timestamp of the most recent direct sampling of option $a_k$, rendering the leapfrog probability $P_{\text{leap}}$ an explicit hazard function that scales monotonically with the temporal interval $(t – \tau_k)$. The parameter $\Delta_{\text{ju\mp}}$ specifies the macroscopic step size of the leapfrog event, calibrated such that $\mu_t(a_k) > \mu_t(a_{\text{active}})$, thereby instantly inverting the optimal policy hierarchy. Because agents operate under a finite temporal horizon $T$, they are subjected to rigorous computational pressure: they must estimate the hazard rate of the leapfrog event while balancing the remaining trial budget against the time required to amortize the informational cost of an exploratory sampling action.
2.2 Algorithmic Implementations: Kalman Filters and Hidden Markov Models
To mathematically capture human decision dynamics during the Leapfrog Task, cognitive modelers employ recursive Bayesian state estimation architectures, predominantly specialized Kalman filters and Hidden Markov Models (HMMs). A canonical Kalman filter assumes that the true reward of an option $k$ at trial $t$ is a hidden state $x_t^{(k)}$ that evolves linearly with Gaussian noise. The agent maintains an internal representation parameterized by an expected mean $\hat{x}_{t|t-1}^{(k)}$ and an associated estimation uncertainty (variance) $\sigma_{t|t-1}^{2(k)}$.
For the actively chosen option, receipt of observed reward $y_t$ prompts an immediate Bayesian update mediated by the Kalman gain $K_t$:
$$K_t^{(k)} = \frac{\sigma_{t|t-1}^{2(k)}}{\sigma_{t|t-1}^{2(k)} + \sigma_{\text{observation}}^2}$$
$$\hat{x}_{t|t}^{(k)} = \hat{x}_{t|t-1}^{(k)} + K_t^{(k)} \left( y_t – \hat{x}_{t|t-1}^{(k)} \right)$$
$$\sigma_{t|t}^{2(k)} = \left( 1 – K_t^{(k)} \right) \sigma_{t|t-1}^{2(k)}$$
For the unchosen, dormant options, no direct observational update occurs ($y_t$ is absent). Consequently, the predictive mean persists while the internal uncertainty variance compounds recursively according to the environmental diffusion parameter $\sigma_{\text{diffusion}}^2$:
$$\sigma_{t+1|t}^{2(k_{\text{dormant}})} = \sigma_{t|t-1}^{2(k_{\text{dormant}})} + \sigma_{\text{diffusion}}^2$$
However, because pure Gaussian Kalman filters assume continuous unimodal drift, they struggle to capture the sharp, non-linear step functions characteristic of discrete leapfrog transitions. To circumvent this limitation, Montague and modern computational psychiatrists deploy hybrid jump-diffusion models or switching state-space architectures implemented via Sequential Monte Carlo (particle filtering) methods or discrete HMM regimes. In an HMM formulation, the latent state space is augmented by a discrete regime indicator $S_t^{(k)} in {0, 1}$, where $S=0$ signifies baseline dormant drift, and $S=1$ indicates that a leapfrog event has occurred. Particle filters maintain a probability distribution over these discrete regime switches by approximating the posterior density through a cloud of weighted particles, enabling the simulated agent to adjust its Kalman gain dynamically in response to abrupt changes in environmental volatility.
2.3 Optimal Policy Benchmarks and Information Theory Metrics
Establishing absolute empirical benchmarks for exploratory performance requires comparing human behavior against theoretical optimal control policies. In stationary and certain constrained non-stationary bandit frameworks, optimal policies can be computed analytically via Gittins indices. A Gittins index maps the complex multidimensional problem of infinite-horizon exploration into a collection of decoupled scalar values, effectively providing a dynamic indifference threshold for each option that natively integrates an exploration bonus proportional to state variance.
However, within the non-linear, finite-horizon architecture of the Leapfrog Task, standard Gittins assumptions break down due to history-dependent hazard rates and non-separable transition dynamics. Consequently, the optimal baseline must be determined via backward induction using dynamic programming over the complete Bayesian belief space. The optimal Bellman value function $V^*(b_t)$ at any information belief state $b_t = (\mathbf{\hat{x}}_t, \mathbf{\sigma}_t^2)$ is defined by:
$$V^*(b_t) = \max_{a in \mathcal{A}} \left{ \mathbb{E}\left[ r(a) mid b_t \right] + \gamma \int \mathcal{P}(b_{t+1} mid b_t, a) V^*(b_{t+1}) db_{t+1} \right}$$
where $\gamma in [0, 1]$ represents the temporal discounting factor. Human deviations from this dynamic programming frontier are rigorously quantified through information-theoretic metrics. Foremost among these is the Shannon entropy $H(b_t)$ of the belief state across all available options, which indexes the global uncertainty of the decision-maker:
$$H(b_t) = -\sum_{k=1}^K p(\text{best}_k mid b_t) \log_2 p(\text{best}_k mid b_t)$$
where $p(\text{best}_k mid b_t)$ represents the posterior probability that option $k$ currently possesses the highest latent payout. Furthermore, the magnitude of internal belief reorganization prompted by an exploratory sampling event is formally captured by the Kullback-Leibler (KL) divergence between the prior and posterior belief distributions: $D_{\text{KL}}(P_{\text{posterior}} parallel P_{\text{prior}})$. By computing cumulative regret—defined as the difference between the maximum achievable reward under an omniscient policy and the realized reward trajectory of the human agent—researchers can precisely locate the computational inefficiencies characterizing different human phenotypes.
3. The Explore-Exploit Trade-Off in Dynamic Non-Stationary Environments
3.1 Cognitive Dilemmas Generated by Leapfrog Mechanics
The architectural genius of Montague’s Leapfrog Task lies in its capacity to generate acute, irresolvable cognitive tension within the decision-maker. Unlike static environments where an agent eventually converges on an optimal choice and remains indefinitely locked in an exploitative state, the Leapfrog Task guarantees that any sustained exploitation will ultimately culminate in suboptimal utility. As an agent exploits an active option, it faces a mounting psychological dilemma: the more stable and rewarding the current choice appears in the short term, the greater the latent probability that an unobserved competitor has leapfrogged into an objectively superior state.
This dynamic introduces substantial cognitive control costs. Deliberately disengaging from a known, immediately rewarding option to sample a historically degraded alternative requires the active suppression of prepotent, reward-seeking motor responses. This disengagement represents a form of cognitive effort, demanding the suspension of dopamine-mediated Pavlovian biases in favor of abstract, model-based counterfactual hypotheses. Human decision-makers experience this cognitive burden as a subjective sense of conflict and tension, which escalates as the unchosen intervals lengthen.
Moreover, the cognitive calculus of the Leapfrog Task is inextricably linked to time-horizon constraints. When an agent possesses a deep temporal horizon (many remaining trials), the epistemic value of exploratory sampling is elevated, as the knowledge gained from discovering a newly elevated leapfrog target can be amortized across numerous subsequent exploitative trials. Conversely, as the remaining trial horizon dwindles toward zero, the expected value of information collapses; the agent lacks the temporal runway required to capitalize on newly acquired knowledge, rendering directed exploration computationally irrational. Human agents frequently display striking vulnerabilities to sequential dependency artifacts and horizon neglect within these trade-offs, failing to dynamically calibrate their exploratory frequency to the diminishing utility horizon.
3.2 Computational Signatures of Strategic Switching
The behavioral execution of an exploratory switch within the Leapfrog Task leaves distinct computational, chronometric, and autonomic footprints. In the domain of chronometrics, reaction time (RT) distributions undergo profound morphological shifts immediately preceding a transition from exploitation to exploration. While stable exploitative sequences are characterized by low-variance, highly accelerated RT profiles indicative of automated motor policies, the trials immediately preceding an exploratory leap display significant RT dilation. This chronometric elongation reflects the accumulation of cognitive control and the recruitment of deliberate deliberative mechanisms required to break out of an automated behavioral attractor.
To mechanistically formalize these dynamics, researchers deploy Drift-Diffusion Models (DDM) and hierarchical race models. In a DDM framework, choice selection is modeled as the continuous accumulation of noisy sensory or value-based evidence toward one of two decision thresholds. In the context of the Leapfrog Task, the drift rate $v$ ceases to be static; it is parametrically modulated on a trial-by-trial basis by the differential between the expected value of the current target and the integrated epistemic value (uncertainty bonus) of the dormant alternatives. As the latent leapfrog hazard increases, the effective drift rate toward the exploitative boundary decreases, causing the decision process to decelerate and increasing the probability that stochastic diffusion will drive the trajectory across the alternative, exploratory threshold.
Concurrently, autonomic markers register these computational shifts in real time. High-speed pupillometry reveals transient baseline pupil dilations and augmented evoked pupil responses prior to exploratory transitions. Mediated by the ascending locus coeruleus-norepinephrine (LC-NE) system, these pupillometric spikes serve as physiological proxies for abrupt reorganizations of network-level cognitive gain. When an internal computational prediction error signals that the current exploitative policy is degraded, a burst of noradrenaline reduces the signal-to-noise ratio of currently dominant cortical representations, effectively “destabilizing” the behavioral attractor and permitting the physical execution of an exploratory switch.
3.3 Heuristics versus Exact Inference in Human Decision-Makers
While the normative mathematical ideal dictates that human agents should execute exact Bayesian filtering combined with backward induction dynamic programming, biological nervous systems operate under strict metabolic and architectural limitations. The human brain is an exemplar of bounded rationality, substituting exact, computationally intractable Bayesian operations with structurally sophisticated heuristics that provide remarkably robust approximations across ecologically valid regimes.
A primary heuristic observed in the Leapfrog Task is a generalized version of the win-stay, lose-shift (WSLS) policy. In classical WSLS, an agent maintains its current action selection if the outcome matches or exceeds expectations, switching exclusively upon receipt of an explicit negative outcome. In the Leapfrog Task, however, agents encounter an environment where an exploited target may yield positive absolute reward even after it has become counterfactually inferior to a leapfrogged alternative. Consequently, human agents implement dynamic threshold-based WSLS heuristics. Rather than reacting to absolute losses, the agent tracks the running reward frequency and implements a “soft” shift heuristic when the observed yield drops below a subjective running average, or when a heuristic counter indexing consecutive exploitative trials crosses a hard-coded internal threshold.
These heuristic shortcuts expose human decision-makers to pervasive cognitive biases. Foremost among these is the Gambler’s Fallacy: after enduring a sequence of trials wherein an unchosen alternative fails to reveal a leapfrog jump upon sampling, human subjects frequently assign an erroneously high probability to an imminent leapfrog event on the subsequent trial, failing to account for the stochastic independence of the underlying hazard function. Furthermore, subjects exhibit pronounced confirmation biases, over-interpreting minor stochastic fluctuations in the exploited option as evidence that the latent alternatives have not yet transitioned. Individual differences in risk aversion (tolerance for outcome variance) and ambiguity aversion (intolerance for missing information) heavily bias these heuristic policies, driving dramatic variations in exploratory strategies across normal and clinical populations.
4. Methodological Constraints and Challenges in Task-Based fMRI
4.1 Temporal Resolution Mismatches and Hemodynamic Blurring
Investigating the neurocomputational mechanics of the Leapfrog Task via functional magnetic resonance imaging forces researchers to confront profound biophysical and temporal limitations. The central methodological paradox of task-based fMRI lies in the temporal discordance between the underlying neurobiology and the proxy signal being captured. The neural computations mediating decision-making—including the continuous recursive evaluation of counterfactual states, the resolution of decision conflict, and the execution of an exploratory motor command—transpire across distributed neuronal assemblies on millisecond timescales (typically within 50 to 400 milliseconds).
Conversely, the BOLD response is not a direct measurement of neural activity, but an indirect, hemodynamically mediated reflection of neurovascular coupling. Following a burst of local neuronal firing, the vascular response—governed by complex cascades of astrocytic signaling, nitric oxide release, and pericyte relaxation—evolves with extreme sluggishness. The canonical hemodynamic response function (HRF) exhibits an initial dip, peaks approximately 5 to 6 seconds following neural onset, and endures an extended post-stimulus undershoot lasting up to 20 to 30 seconds. In rapid event-related experimental designs, where trials occur every 2 to 4 seconds to maintain behavioral engagement, the BOLD responses from consecutive cognitive sub-phases severely overlap.
This severe hemodynamic blurring produces crippling multi-collinearity among linear model regressors. In the Leapfrog Task, each individual trial contains three temporally distinct computational epochs: the evaluation phase (presentation of state cues and retrieval of latent values), the choice phase (action selection and motor execution), and the feedback phase (revelation of reward outcome and computation of prediction errors). If an event-related fMRI design employs static, predictable timing, the design matrix regressors for these sub-phases exhibit near-perfect collinearity, rendering it mathematically impossible to parse whether a prefrontal activation cluster reflects anticipatory value integration, the effort of cognitive disengagement, or feedback-driven prediction error processing. Mitigating this limitation requires aggressive deconvolution strategies, specifically the implementation of Poisson-distributed inter-stimulus intervals (ISI) and trial jittering, which systematically decorrelate the design matrix columns and maximize the statistical efficiency of the general linear model (GLM).
4.2 Susceptibility Artifacts and Signal Dropout in Critical Regions
Perhaps no paradigm exposes the structural vulnerabilities of echo-planar imaging (EPI) as acutely as the Leapfrog Task. The theoretical models advanced by Read Montague identify the ventromedial prefrontal cortex (vmPFC) and the contiguous medial orbitofrontal cortex (mOFC) as the principal computational hubs responsible for integrating latent value, computing common currency signals, and scaling dynamic uncertainty bonuses. Concurrently, midbrain dopaminergic nuclei (the ventral tegmental area and substantia nigra pars compacta) generate the critical reward and state prediction error signals that drive model updating.
Tragically for cognitive neuroscientists, these precise anatomical loci represent the most hostile imaging environments in the human brain. Functional MRI relies overwhelmingly on $T_2^*$-weighted gradient-echo EPI sequences, which are acutely sensitive to macroscopic magnetic field ($B_0$) inhomogeneities. The vmPFC and mOFC sit directly dorsal to the ethmoid sinuses and nasal cavities—anatomical boundaries where human soft tissue interfaces with air. Because air and biological tissue possess radically disparate magnetic susceptibilities ($\chi$), substantial microscopic magnetic field gradients develop across this boundary. These susceptibility-induced gradients accelerate intracellular and extracellular dephasing of transverse magnetization, resulting in devastating signal dropout (complete signal loss) and severe geometric spatial distortions precisely within the ventral and medial orbital surfaces.
To rescue BOLD sensitivity in these computationally non-negotiable regions, researchers must deploy sophisticated methodological adaptations. Standard single-echo gradient-echo sequences must be abandoned in favor of multi-echo EPI or optimized dual-echo acquisition schemes. By acquiring early echo times (e.g., $TE \approx 12-15\text{ ms}$) alongside standard echo times ($TE \approx 30-35\text{ ms}$), early echoes can recover significant transverse magnetization before complete dephasing occurs, allowing for the mathematical recombination of signals across echoes via $T_2^*$ parameter mapping. Furthermore, scanning protocols must implement aggressive slice-angle tilting (typically 30 degrees relative to the anterior commissure-posterior commissure [AC-PC] line) to orient the phase-encoding direction orthogonal to local susceptibility gradients, paired with high-order localized $B_0$ shimming routines tailored specifically to the prefrontal-sinus interface.
4.3 Head Motion and Autonomic Confounds During High-Cognitive Load
A third major methodological hazard in neuroimaging the Leapfrog Task stems from physical head motion and systemic physiological fluctuations induced by high cognitive load. The Leapfrog Task is intentionally structured to generate friction, uncertainty, and cognitive conflict. During periods of sustained exploitation where an agent suspects an impending leapfrog transition, or following an exploratory excursion that tragically reveals a low payout (a failed exploration), human participants experience acute frustration and heightened stress. These psychological shifts translate directly into micro-movements of the head within the RF coil.
In fMRI time-series analysis, head motion represents an insidious confound. Even sub-millimeter displacements (voxel translations $> 0.2\text{ mm}$ or rotations $> 0.2^\text{o}$) generate massive, spurious variance in signal intensity due to spin-history artifacts, wherein biological tissue moves across slice-excitation profiles before transverse magnetization achieves a steady state. More perniciously, these head micro-movements are systematically time-locked to critical computational events: subjects physically flinch or reposition their heads precisely when committing to a high-conflict exploratory switch or upon receiving an unexpected negative outcome. When motion artifacts correlate directly with the experimental design matrix, standard motion realignment parameters entered as nuisance regressors in a GLM fail to fully purge the artifact, yielding rampant false-positive activations that can effortlessly masquerade as prefrontal or insular cognitive control signals.
Simultaneously, the autonomic arousal accompanying exploratory decisions drives profound shifts in systemic physiology. Fluctuations in heart rate and respiratory depth alter arterial carbon dioxide concentration ($Pa\text{CO}_2$), a potent endogenous vasodilator. When a subject holds their breath in anticipation of a high-stakes exploratory feedback, or experiences transient tachycardia during decision conflict, global cerebral blood flow alters independently of local neuronal computation. Failure to monitor and purge these autonomic confounds using dynamic physiological noise correction algorithms—such as RETROICOR (retrospective image-based correction) and respiratory volume per time (RVT) regressors—inevitably contaminates subcortical and frontolimbic BOLD signals, introducing profound systematic error into model-based analyses.
5. Model-Based fMRI: Bridging Computational Latent Variables to BOLD Signals
5.1 Design and Execution of Model-Based Neuroimaging Regressors
The realization of Read Montague’s computational vision demands a radical departure from classical categorical fMRI analysis. In standard subtraction-based neuroimaging, trials are sorted into categorical bins (e.g., “Exploratory Trials” versus “Exploitative Trials”) and contrasted directly via linear subtraction. This coarse methodology completely ignores the continuous, trial-by-trial evolution of cognitive states. An exploratory trial executed when alternative uncertainty is vast and immediate expected value is meager is computationally and neurobiologically distinct from an exploratory trial triggered under low uncertainty by stochastic motor noise.
To overcome this limitation, computational cognitive neuroscience utilizes model-based fMRI. The execution pipeline begins by fitting an algorithmic model—such as a Kalman filter, a jump-diffusion HMM, or a Bayesian active inference architecture—directly to the subject’s behavioral choice string and reaction time distribution. Once the best-fitting free parameters (e.g., learning rates $\alpha$, inverse temperatures $\beta$, information bonus weights $phi$, and hazard rates $lambda$) are identified via maximum likelihood estimation (MLE) or hierarchical Bayesian estimation, the model is run forward in an agent-specific generative mode to extract the continuous, unobservable latent computational variables on a trial-by-trial basis.
These latent time series—including expected value $Q_t(a)$, local prediction error $\delta_t$, Bayesian state uncertainty $\sigma_t^2$, information bonus $I_t$, and decision conflict $C_t$—are mathematically convolved with the canonical hemodynamic response function. The resulting convolved vectors are embedded as parametric modulators within a mass-univariate General Linear Model (GLM). In this framework, the statistical regression equation for an fMRI voxel time series $Y$ is formalized as:
$$Y = \beta_0 + \beta_{\text{onset}} X_{\text{onset}} + \sum_{p=1}^P \beta_p \left( X_{\text{onset}} \cdot M_p \right) + \sum_{n=1}^N \gamma_n Z_n + \epsilon$$
where $X_{\text{onset}}$ represents the unmodulated stick function tracking neural event onsets, $M_p$ denotes the mean-centered vector of the $p$-th latent computational variable, $\beta_p$ indexes the magnitude to which that specific voxel’s BOLD trajectory tracks the computational parameter, $Z_n$ captures nuisance regressors (motion, physiology, drift), and $epsilon$ represents residual Gaussian noise. Crucially, researchers must address orthogonalization dilemmas: because latent variables like expected value and uncertainty frequently correlate naturally, arbitrary Gram-Schmidt orthogonalization within neuroimaging software (such as SPM) can artificially bias parameter estimates, demanding that models be estimated without serial orthogonalization while monitoring variance inflation factors (VIF).
5.2 Statistical Power, Multiple Comparisons, and Random-Effects Modeling
Mapping model-derived latent variables across the entire human brain exposes the analysis to the perilous problem of multiple statistical comparisons. A typical high-resolution whole-brain acquisition matrix contains between 100,000 and 200,000 individual voxels. Conducting mass-univariate hypothesis testing at each voxel simultaneously at an uncorrected alpha level of $\alpha = 0.05$ will inevitably yield 5,000 to 10,000 false-positive voxels purely by chance—an unacceptable vulnerability when tracing subtle computational parametric modulations.
To maintain rigorous family-wise error (FWE) rate control, researchers historically relied on Gaussian Random Field (GRF) theory to execute cluster-extent thresholding. However, following the revelations of Eklund et al. (2016), which demonstrated that parametric cluster-extent methods can produce inflated empirical false-positive rates due to non-Gaussian spatial smoothness structures, the standard in model-based neuroimaging has transitioned toward non-parametric permutation testing. By permuting behavioral regressor trajectories relative to observed fMRI time series across thousands of iterations, investigators empirically generate null distributions of cluster mass, safeguarding against spurious activations in complex cortical surfaces.
At the group level, statistical inference must be executed via Linear Mixed-Effects (LME) models or hierarchical random-effects (RFX) frameworks. Human participants do not share identical neural gain functions; variance in biological parameters such as dopamine receptor availability or prefrontal cortical thickness results in pronounced inter-subject variability in the slope ($\beta_p$) of the BOLD-computational relationship. RFX models treat individual subject parameter estimates as random draws from an underlying population distribution, ensuring that statistical inferences generalize legitimately to the human population. Furthermore, when defining regions-of-interest (ROIs) for in-depth computational dissection, researchers must strictly enforce non-circular inference criteria (mitigating “double dipping”), utilizing anatomically defined masks or independent functional localizers to guarantee statistical integrity.
5.3 Evaluating Model Fit: Bayesian Model Selection in Neuroimaging
A recurring hazard in computational neuroimaging is the commitment to a single algorithmic architecture without demonstrating that the hypothesized model provides the most parsimonious and accurate account of both behavioral performance and neural variance. A researcher might assume that subjects solve the Leapfrog Task using a fully realized Bayesian Kalman filter, yet the empirical data might be vastly better explained by a dynamic heuristic model or an asymmetric Temporal Difference Reinforcement Learning (TDRL) algorithm.
To scientifically validate computational architectures, cognitive neuroscientists deploy formal Bayesian Model Selection (BMS). BMS operates across a competitive model space comprising multiple competing computational hypotheses: for example, comparing a Model-Free RL architecture, an Uncertainty-Augmented RL model (Model-Free + UCB), an exact Bayesian Jump-Diffusion Kalman filter, and an empirical Win-Stay Lose-Shift heuristic. For each subject, the marginal likelihood (model evidence) of each candidate model is computed, penalizing for model complexity via metrics such as the Bayesian Information Criterion (BIC) or Laplace approximations of the free energy:
$$\mathcal{F} \approx \ln p(D mid \hat{\theta}, m) – \frac{1}{2} \text{dim}(\theta) \ln(N)$$
These individual log-evidences are then integrated into a random-effects BMS framework at the group level, which treats the choice of computational model as a multinomial latent variable across the population. This generates the Exceedance Probability ($EP$) and the Protected Exceedance Probability ($PEP$), indexing the statistical probability that a given computational model is more prevalent in the population than any other tested alternative, accounting for the probability that all differences could have arisen by chance. Once the winning computational model is identified, its specific latent parameters are demonstrated to account for unique, functionally specific BOLD variance across targeted neural networks, establishing a verified mapping between algorithmic computation and biological wetware.
6. Prefrontal Cortical Topography: Medial vs. Lateral Control Networks
6.1 The Role of the Ventromedial Prefrontal Cortex (vmPFC) in Value Integration
At the core of the frontal valuation network sits the ventromedial prefrontal cortex, encompassing medial areas of Brodmann areas 10, 11, 14, and 32. In the computational anatomy of the Leapfrog Task, the vmPFC serves as the central biophysical hub for value integration, executing the translation of multidimensional choice attributes into a unified, common currency value metric. Whenever an agent engages in the evaluation of a candidate target, the local BOLD signal within the vmPFC scales linearly with the subjective expected value (EV) of that option.
Under exploitative stability, where an agent continuously harvests reward from an established active target, the vmPFC demonstrates robust, sustained activation. This activation profile does not merely encode the nominal magnitude of the reward; it dynamically tracks the integrated expected utility, discounted continuously by the agent’s internal estimation of risk and mounting environmental uncertainty. As consecutive exploitative trials elapse without sampling the alternatives, the vmPFC BOLD signal exhibits a characteristic decay. This attenuation reflects the cognitive discounting of the incumbent option’s subjective utility: even though the active target continues to yield nominal payouts, its subjective value is continuously eroded by the escalating opportunity cost dictated by the unobserved leapfrog hazard.
The computational necessity of the vmPFC in this capacity has been rigorously confirmed by clinical lesion studies and localized pharmacological inactivation models. Patients suffering from focal bilateral vmPFC damage exhibit catastrophic deficits when navigating non-stationary value environments. Rather than maintaining a coherent, value-maximizing policy punctuated by strategically timed directed exploration, vmPFC-lesioned individuals display chaotic behavioral fragmentation. They either exhibit pathological perseveration on options that have long since dropped below counterfactual thresholds, or they degenerate into hyper-frequent, unstructured behavioral switching. Lacking an intact neural substrate to represent the common currency value metric, their cognitive control architecture fails to compute the differential between immediate utility and counterfactual opportunity costs, rendering deliberate value-based exploration impossible.
6.2 Frontopolar Cortex (BA 10) and the Arbitration of Alternative Goals
While the medial prefrontal wall tracks the absolute utility of current choices, the most rostral territory of the frontal lobes—the frontopolar cortex (BA 10), situated in the lateral and polar aspects of the prefrontal architecture—executes a computational operation uniquely critical to the Leapfrog Task. Pioneering neuroimaging investigations by Daw et al. (2006) and Boorman et al. (2009) demonstrated that the lateral frontopolar cortex does not track the value of the currently selected option; instead, it specifically and continuously monitors the relative counterfactual value of the best unchosen alternative.
In the Leapfrog Task, as an agent remains engaged with the active target, the latent probability distribution of the dormant options expands. The lateral frontopolar cortex actively maintains a running representation of the counterfactual state space, tracking the evidence that a dormant alternative has leapfrogged the incumbent target. The BOLD signal in BA 10 scales directly with the magnitude of unchosen option uncertainty and the integrated epistemic value bonus. When the counterfactual value representation maintained within BA 10 surpasses the subjective value signal maintained in the vmPFC, the frontopolar cortex initiates a top-down executive command, signaling an imperative to divert behavioral resources away from the ongoing task.
This frontopolar monitoring architecture operates within a defined hierarchical network. Lateral BA 10 does not directly innervate primary motor effectors; rather, it coordinates with the dorsal anterior cingulate cortex and the pre-supplementary motor area (pre-SMA). When the epistemic threshold is crossed, BA 10 sends excitatory efferent projections to the pre-SMA, actively reconfiguring the motor policy from an automated exploitative key-press to a deliberative, exploratory sampling response directed toward the counterfactual target. Thus, the frontopolar cortex functions as the brain’s computational arbiter of alternative goals, allowing human agents to hold alternative behavioral trajectories in a suspended, active monitoring state while executing an ongoing baseline policy.
6.3 Dorsolateral Prefrontal Cortex (dlPFC) and Working Memory Updating
Flanking the frontopolar and medial networks, the dorsolateral prefrontal cortex (dlPFC; BA 9/46) acts as the operational workhorse for executive cognitive control and working memory maintenance throughout the Leapfrog Task. Tracking leapfrog dynamics imposes severe working memory demands: an agent cannot reliably estimate hazard rates or particle filter weights without retaining a precise historical trace of previous choice outcomes, elapsed trials since last sampling ($\tau_k$), and the local volatility of the environment.
The dlPFC implements this computational requirement via sustained, persistent recurrent neuronal firing. Electrophysiological and high-resolution neuroimaging studies indicate that dlPFC microcircuits maintain the parametric representations of unchosen option reward trajectories across extended temporal intervals. During the inter-trial intervals of the Leapfrog Task, when the screen is blank or displaying a fixation cross, the dlPFC exhibits persistent BOLD elevation, indexing the energetic maintenance of the generative model’s state variables within local prefrontal recurrent networks.
Furthermore, the dlPFC is heavily recruited to enforce top-down cognitive inhibition over automated, habitual behavior. When an agent resolves to execute an exploratory choice, it must deliberately suppress the powerful motor habit that has crystallized over multiple consecutive exploitative trials. The dlPFC, operating in tandem with the posterior parietal cortex within the canonical frontoparietal cognitive control network, sends dense glutamatergic projections to the basal ganglia, providing the top-down bias required to override prepotent motor execution and re-route behavioral output toward the unchosen target.
7. Cingulate Cortex Dynamics: Monitoring Conflict, Volatility, and Surprise
7.1 Dorsal Anterior Cingulate Cortex (dACC) and Strategic Shifting
Deep within the medial longitudinal fissure, the dorsal anterior cingulate cortex (dACC; BA 24/32) plays a quintessential computational role in detecting internal friction and signaling the need to abandon established behavioral policies. In the theoretical framework of the Leapfrog Task, the dACC functions as a continuous monitor of decision conflict. Decision conflict escalates precisely when the expected utility of the current exploitative option and the counterfactual expected utility of an unobserved leapfrog alternative approach parity.
When an agent is deeply confident in the superiority of its current target, conflict is minimal, and dACC BOLD activity remains quiescent. However, as the latent hazard function indicates that a leapfrog event is increasingly probable, the decision landscape becomes ambiguous. The dACC responds to this equilocal state with an escalating, non-linear surge in BOLD amplitude. According to the Expected Value of Control (EVC) theory advanced by Shenhav, Botvinick, and Cohen, the dACC integrates this conflict signal along with the estimated costs of physical or mental exertion, computing whether it is computationally worthwhile to allocate the intense cognitive control required to interrupt the ongoing policy.
Crucially, the dACC does not merely register static conflict; it acts as a primary computational engine for tracking environmental volatility. Work by Behrens et al. (2007) established that the dACC BOLD response parametrically tracks the variance of the generative model’s transition distribution. When the leapfrog step sizes ($\Delta_{\text{ju\mp}}$) become erratic or the jump probability ($P_{\text{leap}}$) shifts unpredictably, the dACC scales its output, directly modulating downstream learning rates ($\alpha$) across the broader prefrontal-striatal axis. In doing so, the dACC serves as the biological trigger that mandates a policy shift, instructing the motor architecture to cease exploitation and initiate exploratory foraging.
7.2 Divergent Roles of Anterior Mid-Cingulate and Posterior Cingulate Cortices
While the dACC orchestrates the immediate response to decision conflict and volatility, contiguous and posterior cingulate regions demonstrate specialized functional divergence during the Leapfrog Task. The anterior mid-cingulate cortex (aMCC) is specifically recruited when an exploratory action yields an unexpected outcome. If an agent executes an exploratory choice expecting to find a leapfrogged high-yield target, but instead uncovers an unexpectedly degraded payout, the aMCC fires intensely. This region encodes the negative discrepancy between anticipated and realized information gain, translating computational surprise into visceral motor-readiness adjustments.
Simultaneously, the posterior cingulate cortex (PCC; BA 23/31), a foundational node of the Default Mode Network (DMN), exhibits dynamic network-level switching. During periods of sustained, highly focused exploitation, the PCC is profoundly deactivated, reflecting the suppression of task-irrelevant, self-referential cognition in favor of the externally directed frontoparietal control network. However, during the precise temporal window wherein an agent decides to transition from exploitation to exploration, the PCC experiences a sharp, transient BOLD rebound.
This transient PCC recruitment reflects a profound shift across large-scale cortical states. The PCC works in synchrony with the salience network to facilitate the de-differentiation of cognitive focus, permitting the agent to retrieve broad contextual memories and re-evaluate its global behavioral strategy. Furthermore, during the feedback phase of the Leapfrog Task, the PCC specifically encodes counterfactual regret. When the outcome screen reveals that an unchosen option has indeed leapfrogged to an astronomical value while the agent remained stubbornly anchored to a sub-optimal choice, the magnitude of the counterfactual regret signal ($r_{\text{unchosen}} – r_{\text{chosen}}$) correlates robustly with the amplitude of the PCC BOLD response, driving retrospective belief revision.
7.3 Computational Prediction Errors in the Cingulate Node
A critical neurocomputational distinction enforced within Montague’s investigative paradigms is the operational divergence between standard reward prediction errors (RPEs) and state prediction errors (SPEs). While classical RPEs quantify the difference between received reward and expected reward ($\delta_{\text{RPE}} = r_t – \hat{V}_t$), state prediction errors capture deviations in the abstract, non-scalar structural transitions of the environment ($\delta_{\text{SPE}} = s_{t+1} – \hat{s}_{t+1}$), completely independent of monetary or primary utility.
The cingulate cortex, particularly the pre-cingulate and rostral cingulate zones, provides a dedicated computational substrate for tracking these unsigned state prediction errors (often conceptualized as pure perceptual or structural “surprise”). When an agent samples an alternative in the Leapfrog Task and discovers that a leapfrog event has occurred, the cingulate fires intensely even if the absolute monetary reward obtained on that individual trial is modest. The cingulate computes the magnitude of the model update required to realign the internal Bayesian belief distribution with the newly observed environmental state.
This neuroimaging profile aligns precisely with decades of human electrophysiological investigations. The classic Error-Related Negativity (ERN) and Feedback-Related Negativity (FRN)—event-related potentials (ERPs) peaking between 100 and 300 milliseconds post-feedback—have their primary dipolar electrical generators localized directly within the anterior cingulate cortex. By combining high-density electroencephalography with simultaneous model-based fMRI, researchers have confirmed that the trial-by-trial amplitude of the FRN correlates with the precision-weighted state prediction error convolved into the cingulate BOLD signal, illustrating how the cingulate dynamically scales learning rates in real time to restructure exploratory policies.
8. Subcortical and Dopaminergic Signatures in the Leapfrog Landscape
8.1 Striatal Dissociations: Ventral versus Dorsal Striatum
Deep beneath the cerebral mantle, the basal ganglia provide the operational subcortical machinery that executes value learning and policy gating. In the computational topography of the Leapfrog Task, functional neuroimaging delineates an essential, double-dissociation between the ventral striatum (encompassing the nucleus accumbens and ventral putamen) and the dorsal striatum (comprising the dorsal caudate nucleus and dorsal putamen).
The ventral striatum (VS) functions as the primary subcortical engine for reward prediction error processing. Whenever the feedback screen confirms the outcome of a chosen option, the VS BOLD signal tracks the signed parametric RPE with extraordinary fidelity. Under exploitative stability, this VS signal exhibits gradual extinction as the target’s payout becomes fully predicted by antecedent visual cues. However, when an agent executes an exploratory choice and uncovers a massive, newly leapfrogged reward, the VS responds with a massive, hyper-elevated BOLD surge. Concurrently, computational modeling reveals that the VS also tracks the epistemic value bonus: when an agent is presented with an option characterized by high Bayesian variance, the anticipatory cue-locked VS BOLD signal is elevated, indicating that the human striatum directly treats uncertainty reduction as an intrinsically rewarding property.
In contrast, the dorsal striatum is functionally dedicated to the encoding of action-value mappings ($Q(s, a)$) and the structural consolidation of behavioral policies. The dorsal caudate nucleus is heavily engaged during the early, fluid phases of the Leapfrog Task, actively interacting with the dorsolateral prefrontal cortex to maintain flexible model-based goal representations. As an agent identifies a superior leapfrog target and transitions into an extended exploitative sequence, computational control progressively shifts from the caudate to the dorsal putamen. The putamen mediates the consolidation of action-selection into automated, model-free motor routines. During this phase, dorsal putamen activation stabilizes, reflecting the execution of crystalline, low-effort exploitative loops that conserve prefrontal cognitive resources until a cingulate-driven conflict signal disrupts the cycle.
8.2 Midbrain Dopamine Systems: SNc and VTA Contributions
The computational operations of both the prefrontal cortex and the striatum are fundamentally driven by ascending monoaminergic projections originating within the midbrain: specifically, the substantia nigra pars compacta (SNc) and the ventral tegmental area (VTA). Understanding the role of dopamine in the Leapfrog Task requires parsing the critical operational dichotomy between phasic and tonic dopaminergic signaling profiles.
Phasic dopamine release—characterized by rapid, transient (100–500 millisecond) burst firing of midbrain neurons—represents the biological instantiation of Montague, Dayan, and Schultz’s reward prediction error algorithm. When an agent samples a leapfrogged alternative and receives an outcome that shatters prior baseline expectations, a massive phasic burst of dopamine is propagated throughout the striatal and prefrontal terminal fields. This phasic signal acts as a potent synaptic plasticity gating mechanism: it induces long-term potentiation (LTP) at active cortical-striatal synapses within the direct pathway, stamping in the newly discovered leapfrog target as the new preferred attractor state.
Conversely, tonic dopamine levels—the slow, ambient, baseline concentrations of extracellular dopamine fluctuating across minutes—play an entirely distinct computational role: the dynamic regulation of the explore-exploit trade-off. Groundbreaking theoretical frameworks propose that tonic dopamine sets the agent’s internal “opportunity cost of time” and scales response vigour. Within the Leapfrog Task, when tonic dopamine tone is high, the subjective cost of latency is elevated, promoting rapid exploitation of known resources. Conversely, when tonic dopamine availability declines or is pharmacologically suppressed, the relative cost of exploration diminishes, lowering the activation barriers required to trigger behavioral switches. High-resolution neuromelanin-sensitive and iron-sensitive fMRI at ultra-high fields has successfully isolated these tiny midbrain nuclei, demonstrating that baseline BOLD shifts within the VTA parametrically predict individual differences in baseline exploratory tendencies.
8.3 The Basal Ganglia Indirect Pathway in Behavioral Inhibition
Executing an exploratory choice in the Leapfrog Task requires more than just the abstract identification of a counterfactual opportunity; it physically mandates the absolute motoric arrest of the ongoing exploitative response. The biological architecture responsible for this emergency braking mechanism is the indirect and hyperdirect pathways of the basal ganglia, anchored by the subthalamic nucleus (STN).
When the dACC and frontopolar cortex detect high decision conflict and resolve to explore, they do not merely excite alternative motor representations. Instead, they exploit the hyperdirect pathway, sending rapid glutamatergic projections directly to the STN. The STN acts as a global, non-specific “brake” on the motor system. By delivering powerful, widespread excitation to the internal segment of the globus pallidus (GPi) and the substantia nigra pars reticulata (SNr), the STN drives intense GABAergic inhibition over the ventral thalamus. This transiently freezes all motor output, abruptly elevating the decision boundary within the drift-diffusion framework.
This STN-mediated pause is computationally indispensable: it prevents the premature, automatic execution of the prepotent exploitative habit, buying the slow, deliberative prefrontal networks the necessary hundred milliseconds required to compute counterfactual updates and assemble the exploratory motor command. Once the exploratory policy is formally selected by the dlPFC, focused striatal projections via the direct pathway specifically inhibit the corresponding pallidal sub-territory, selectively releasing the brake on the targeted exploratory action. Functional MRI studies tracking high-conflict exploratory switches reliably capture robust parametric activations within the STN and the external globus pallidus (GPe), validating the mechanistic role of the indirect basal ganglia architecture in enabling strategic cognitive diversion.
9. Computational Connectivity: Multivariate and Network-Level Analyses
9.1 Psychophysiological Interactions (PPI) and Generalized PPI (gPPI)
While mass-univariate parametric mapping reveals the isolated functional nodes recruited during the Leapfrog Task, it fails to capture the dynamic, inter-regional information flow that coordinates exploratory behavior. Cognitive computation is inherently a distributed, network-level phenomenon. To elucidate how inter-regional communication is dynamically modulated by task demands, researchers deploy generalized Psychophysiological Interaction (gPPI) analyses.
In a gPPI architecture, the statistical model tests for specific, task-dependent changes in effective connectivity between an a priori seed region and the remainder of the brain. The regression equation incorporates three critical vectors: the extracted physiological time series of the seed region (deconvolved into an estimate of underlying neuronal activity), the psychological task regressor (e.g., Exploratory Choice Trials versus Exploitative Baseline Trials), and the psychophysiological interaction term (the mathematical product of the deconvolved neural signal and the psychological regressor). In the Leapfrog Task, placing a seed in the lateral frontopolar cortex (BA 10) reveals a striking functional reorganization during exploratory transitions: BA 10 exhibits a dramatic surge in positive effective connectivity with the ventral striatum and the pre-supplementary motor area precisely during the trial onsets of directed exploration.
Concurrently, gPPI analyses seeded in the dorsal anterior cingulate cortex illuminate top-down regulatory dynamics. As decision conflict escalates prior to an exploratory leap, the dACC establishes intensified functional couplings with the dorsolateral prefrontal cortex and the subthalamic nucleus. This directional connectivity reflects the top-down mobilization of cognitive control networks to enforce the motor pause and allocate attentional resources toward counterfactual evaluation. However, researchers must navigate significant methodological constraints when interpreting PPI results: the standard deconvolution of the hemodynamic response function relies on deterministic assumptions that vary widely across disparate cortical and subcortical vascular beds, demanding cautious, multi-seed validation.
9.2 Multi-Voxel Pattern Analysis (MVPA) and Representational Similarity Analysis (RSA)
Mass-univariate fMRI is fundamentally constrained by spatial resolution: an individual imaging voxel contains hundreds of thousands of neurons. If a cortical region contains intermingled, highly distributed neural subpopulations that separately encode the value of Option A and Option B, the average spatial BOLD signal across the voxel may display zero net change during choice evaluation. To penetrate this spatial barrier, computational neuroscientists utilize multivariate machine learning frameworks, primarily Multi-Voxel Pattern Analysis (MVPA) and Representational Similarity Analysis (RSA).
Rather than averaging voxel intensities within an ROI, MVPA treats the fine-grained spatial topography of BOLD amplitudes across a cluster of voxels as a high-dimensional vector. Using linear support vector machines (SVMs) or regularized logistic regression, researchers can train decoders to classify the cognitive state of the participant. In the Leapfrog Task, cross-validated MVPA decoders trained on distributed patterns within the frontopolar cortex can successfully decode the continuous, latent belief state of the agent—predicting which unchosen alternative the agent believes has leapfrogged several trials before the subject physically commits to an exploratory choice. This provides irrefutable empirical evidence that the frontopolar cortex maintains persistent, high-fidelity traces of counterfactual alternatives in a forward-looking, anticipatory format.
Representational Similarity Analysis takes this methodology a step further by constructing Representational Dissimilarity Matrices (RDMs). An RDM computes the pairwise geometrical distance (typically via correlation distance, $1 – r$) between the multi-voxel neural patterns evoked by every experimental trial condition. Concurrently, computational modelers construct model-derived RDMs representing the theoretical distances predicted by competing algorithmic frameworks (e.g., exact Bayesian entropy versus simple trial-count heuristics). By calculating the Spearman rank correlation between the neural RDMs and the computational RDMs, researchers can quantitatively determine where in the brain specific computational geometries are physically instantiated. RSA studies of the Leapfrog Task reveal that while the vmPFC pattern geometry matches an integrated scalar utility model, the pattern geometry of the lateral frontopolar and intraparietal cortices uniquely reflects the multidimensional geometry of Bayesian state uncertainty.
9.3 Dynamic Causal Modeling (DCM) of Frontostriatal Networks
To move beyond correlation and statistical interaction toward explicit, directed biophysical causality, neuroimaging methodologists implement Dynamic Causal Modeling (DCM). DCM treats the brain as a non-linear dynamic system wherein hidden neuronal states generate observed hemodynamic responses through biophysically realistic models of neurovascular coupling. In DCM, researchers construct a defined anatomical network architecture—typically comprising the vmPFC, frontopolar cortex, dACC, and ventral striatum—and formulate specific mathematical differential equations governing directed neuronal influences:
$$\frac{dz}{dt} = \left( A + \sum_j u_j B^{(j)} \right) z + C u$$
In this formulation, the matrix $A$ represents the endogenous, baseline directed connectivity between brain regions in the absence of task inputs; the matrix $C$ encodes the direct driving inputs of experimental stimuli (e.g., presentation of the Leapfrog task display); and the critical matrix $B^{(j)}$ specifies how specific computational conditions (such as the computational demand to explore versus exploit) modulate the strength of specific directed connections.
Using Bayesian model reduction and parameter estimation over empirical fMRI data acquired during the Leapfrog Task, DCM reveals that exploratory switching is universally mediated by an asymmetric top-down modulation. When an agent transitions into exploration, the directed inhibitory connection from the dACC to the exploitative value-tracking node of the vmPFC strengthens, while the directed excitatory connection from the lateral frontopolar cortex to the ventral striatal gating mechanism is dramatically amplified. This causal modeling confirms the long-standing theoretical premise of Montague’s neuroeconomic architecture: exploratory choice selection is not a passive bottom-up sensory-driven phenomenon, but an active, top-down prefrontal reconfiguration that systematically overrides striatal associative habits.
10. Inter-Subject Variability, Clinical Phenotypes, and Psychopathology
10.1 Dimensional Alterations in Substance Use Disorders and Addiction
Read Montague’s foundational contributions to computational neurobiology were never intended to remain confined to basic cognitive taxonomy; their ultimate destiny lies in the revolutionizing of clinical psychiatry. Under the framework of computational psychiatry, mental health disorders are conceptualized not as discrete, categorical diagnostic checklist entities, but as systematic breakdowns in the latent parameters of underlying computational algorithms. In this paradigm, the Leapfrog Task serves as a diagnostic computational assay of exceptional sensitivity.
When applied to individuals suffering from substance use disorders (SUDs)—including severe alcohol, stimulant, and opioid dependencies—the Leapfrog Task exposes profound, dimensional computational pathology. The hallmark phenotype of addiction is compulsive perseveration: patients remain pathologically anchored to an exploitative target long after environmental dynamics have driven its utility into severe decline. In the Leapfrog Task, individuals with substance dependence demonstrate a striking blunting of their sensitivity to epistemic value. While healthy controls deploy directed exploration precisely as Bayesian uncertainty bounds expand, addicted individuals display near-complete suppression of the information bonus parameter ($\phi \approx 0$).
Model-based fMRI illuminates the biological architecture driving this failure. In SUD cohorts, the ventral striatum displays severely blunted prediction error signaling when encountering unexpected outcomes during exploratory sampling. Concurrently, functional and structural connectivity within the frontostriatal tracts connecting the dlPFC and frontopolar cortex to the nucleus accumbens is profoundly degraded. The addict’s brain fails to compute the opportunity cost of continued exploitation, driven by chronic, drug-induced neuroadaptations in striatal $D_2$ receptor density and the systematic decoupling of prefrontal cognitive control networks, leaving the individual trapped in an automated, maladaptive behavioral attractor.
10.2 Mood Disorders: Anhedonia, Major Depression, and Apathy
Conversely, the computational pathology characterizing major depressive disorder (MDD) and clinical anhedonia manifests through an entirely distinct parameter space within the Leapfrog architecture. A core debilitating symptom of clinical depression is a profound apathy and an unwillingness to exert cognitive effort. In the context of the Leapfrog Task, directed exploration is computationally expensive: it requires the recruitment of frontopolar working memory networks, the active suspension of automated motor output, and the confrontation of ambiguous, potentially low-reward outcomes.
When depressed patients navigate the Leapfrog Task, their behavioral profiles reveal marked distortions in learning rates and asymmetric prediction error updating. Classical reinforcement learning models fitted to depressive cohorts reveal an inflated learning rate for negative prediction errors ($\alpha^-$) coupled with a severely attenuated learning rate for positive outcomes ($\alpha^+$). When an exploratory action yields a sub-optimal reward, the depressed agent dramatically over-weights this negative feedback, rapidly abandoning the exploratory trajectory and retreating into passive, low-yield exploitation. Furthermore, patients suffering from profound anhedonia display a baseline collapse in their willingness to exert the cognitive effort required to compute counterfactual trajectories, preferring to accept deteriorating local rewards rather than engage the prefrontal metabolic machinery.
Functional neuroimaging correlates of this depressive phenotype demonstrate persistent hypoactivity within the ventromedial prefrontal cortex and the ventral striatal reward hubs. During counterfactual feedback, when an unchosen option is revealed to have leapfrogged, depressed individuals exhibit blunted BOLD responses in the frontopolar cortex, reflecting a failure of the counterfactual monitoring apparatus. Concurrently, their generative models exhibit an overestimation of environmental volatility, generating a chaotic, unstructured form of random exploration—driven by despair rather than epistemic curiosity—that further destabilizes real-world psychosocial functioning.
10.3 Schizophrenia and Aberrant Salience Processing
In schizophrenia, the computational architecture of the Leapfrog Task completely disintegrates, providing a remarkable experimental model for the phenomenon of aberrant salience. The dopamine hypothesis of schizophrenia posits that unprovoked, chaotic phasic dopamine release occurs independently of environmental stimuli, stamping neutral or irrelevant sensory events with false computational significance.
When navigating the Leapfrog Task, individuals with schizophrenia exhibit hyper-frequent, completely disorganized exploratory switching. Unlike healthy controls whose exploratory frequency is tightly regulated by Bayesian uncertainty bounds and hazard rates, schizophrenic patients switch away from highly rewarding exploitative targets after single, completely inconsequential trials. Computational modeling reveals that their internal generative models fail to maintain stable precision-weighted prediction errors. Because their midbrain dopaminergic system fires stochastically, minor baseline noise in the exploited option’s payout ($\epsilon_{\text{exploit}}$) is processed as a massive, world-altering prediction error, falsely signaling that an unobserved option has leapfrogged.
Functional neuroimaging in this population reveals severe computational decoupling across the prefrontal-cingulate-striatal network. The dorsal anterior cingulate cortex exhibits continuous, hyper-active BOLD elevations, signaling non-existent conflict and ungrounded environmental volatility. Concurrently, the frontopolar cortex fails to display the organized, parametric tracking of counterfactual state distributions; instead, multi-voxel pattern classifiers fail to decode coherent belief trajectories from prefrontal patterns. The brain of the schizophrenic patient is trapped in a computational state of permanent, unstructured epistemic panic, incessantly destabilizing its behavioral policies in response to internally generated neural noise.
11. Technological Innovations and Advanced Neuroimaging Approaches
11.1 Ultra-High Field (7 Tesla) fMRI Applications in Subcortical Imaging
To overcome the spatial and biophysical limitations that historically hampered 3 Tesla neuroimaging of the Leapfrog Task, cutting-edge laboratories are transitioning to ultra-high field (UHF) 7 Tesla fMRI systems. The physical jump from 3T to 7T provides an enormous boost in the signal-to-noise ratio (SNR) and functional contrast-to-noise ratio (CNR), scaling approximately supra-linearly with static magnetic field strength ($B_0$).
In the context of the Leapfrog Task, 7T imaging delivers two revolutionary methodological capabilities. First, it enables sub-millimeter spatial resolution (e.g., $0.7-0.8\text{ mm}$ isotropic voxels), allowing researchers to move beyond coarse regional activations toward layer-specific (laminar) functional mapping within the prefrontal cortex. Computational models predict that the top-down cognitive signals that drive exploration (originating in frontopolar and dlPFC networks) should terminate primarily within the deep infragranular layers (Layers V and VI) of target motor and striatal structures, while bottom-up sensory feedback and prediction error updates enter via the granular Layer IV. With laminar 7T fMRI, investigators can physically test these computational circuit predictions, dissociating feedforward from feedback informational streams during explore-exploit switches.
Second, 7T imaging provides the precise spatial resolution and magnetic contrast required to definitively isolate the subcortical and midbrain nuclei that orchestrate exploration. At 3 Tesla, the substantia nigra pars compacta and the ventral tegmental area blur into an indivisible midbrain mass, severely confounded by pulsatile cerebrospinal fluid (CSF) motion. At 7 Tesla, using specialized high-resolution anatomical sequences (such as 3D Magnetization Prepared 2 Rapid Acquisition Gradient Echoes [MP2RAGE]) combined with localized zoomed EPI acquisitions, researchers can resolve the internal architecture of the VTA, the SNc, the subthalamic nucleus, and the locus coeruleus. This allows for the simultaneous, segregated mapping of noradrenergic exploration bonuses and dopaminergic prediction errors in the living human brain with unprecedented fidelity.
11.2 Simultaneous Multimodal Imaging: fMRI-EEG and Pharmacological Probes
While 7T imaging solves spatial resolution limitations, it cannot alter the sluggish biophysical physics of the BOLD hemodynamic response function. To capture the full spatiotemporal trajectory of exploratory decision-making, researchers deploy simultaneous multimodal fMRI-EEG recording systems. By placing high-density, MR-compatible EEG caps (64 to 128 electrodes) on participants within the bore of the magnet, scientists achieve sub-millisecond temporal resolution perfectly locked to millimeter-scale spatial localization.
This multimodal integration is uniquely powerful when tracing cingulate and frontal dynamics during the Leapfrog Task. The millisecond precision of electrophysiology captures mid-frontal theta-band oscillations (4–8 Hz), an oscillatory signature long hypothesized to serve as the brain’s universal clock for cognitive control and conflict resolution. Simultaneous recording demonstrates that bursts of mid-frontal theta power precede exploratory transitions by exactly 250 to 350 milliseconds. By using single-trial EEG theta power as a continuous parametric modulator in the concurrent fMRI GLM, researchers have definitively proven that these fleeting oscillatory bursts originate directly within the dACC and pre-SMA, instantly triggering the downstream recruitment of the subthalamic nucleus to pause ongoing behavior.
Furthermore, to establish unambiguous causal links between computational parameters and neurotransmitter systems, multimodal designs incorporate rigorous pharmacological challenges. By administering double-blind, placebo-controlled pharmacological probes—such as the dopamine precursor L-DOPA, the dopamine $D_2/D_3$ receptor antagonist sulpiride, or the noradrenergic beta-blocker propranolol—investigators can directly manipulate the neurobiological substrate. Pharmacological computational modeling reveals that elevating dopamine via L-DOPA systematically attenuates directed exploration bonuses ($phi$), shifting human agents toward aggressive exploitation, whereas blocking noradrenaline degrades the precision of hazard-rate estimation, validating the theoretical tenets of Montague’s neuroeconomic architecture.
11.3 Machine Learning and Artificial Neural Network Comparison Architectures
The contemporary frontier of computational neurobiology involves benchmarking human neural dynamics against deep artificial neural networks (ANNs) and Deep Reinforcement Learning (DRL) agents trained on identical task architectures. When an artificial recurrent neural network (RNN), augmented with long short-term memory (LSTM) units or gated recurrent units (GRUs), is trained to solve the Leapfrog Task via policy gradient algorithms, it spontaneously develops internal representations that mirror the biological brain.
By extracting the high-dimensional hidden unit activation vectors of these in silico agents across thousands of simulated task trials, researchers construct artificial representational spaces. Using Canonical Correlation Analysis (CCA) and Representational Similarity Analysis, these artificial network activations can be compared directly to human fMRI BOLD patterns recorded during the identical task epochs. Strikingly, deep reinforcement learning agents that successfully master the Leapfrog Task spontaneously evolve specialized internal nodes that track counterfactual unchosen value and Bayesian entropy, displaying geometric trajectories mathematically congruent to the activity patterns recorded within human frontopolar cortex (BA 10) and dorsal anterior cingulate cortex.
Moreover, researchers deploy Generative Adversarial Frameworks (GANs) and variational autoencoders (VAEs) to synthetically simulate fMRI BOLD patterns. By conditioning these deep generative models on specific computational latent variables (e.g., conditioning on a state prediction error $\delta_{\text{SPE}} = 2.5$), the generative network can synthesize realistic, whole-brain spatial BOLD patterns. Comparing the synthesized volumes against empirical human scans enables researchers to stress-test the completeness of their computational models, determining whether the hypothesized algorithmic parameters account for the totality of distributed neural variance or whether critical computational dimensions remain undiscovered.
12. Future Trajectories: Multi-Scale Computational Neuroeconomics
12.1 Intracranial Recordings and Single-Unit Validation in Humans
The definitive frontier in validating Read Montague’s computational neurobiology lies in crossing the translational bridge from non-invasive macroscopic imaging to direct, intracranial cellular electrophysiology in human patients. While fMRI captures the blood-oxygenation changes of hundreds of thousands of neurons simultaneously, neurosurgical interventions for intractable epilepsy provide an extraordinary, rare scientific window into microscopic computational circuits.
Neurosurgical patients undergoing pre-surgical clinical evaluation are frequently implanted with invasive intracranial depth electrodes, a procedure known as stereo-electroencephalography (sEEG), or subdural electrocorticographic (ECoG) grids. When these patients volunteer to perform the Leapfrog Task while resting in hospital beds, researchers can record local field potentials (LFPs) and single-neuron action potentials directly from within the human ventromedial prefrontal cortex, the anterior cingulate cortex, the amygdala, and the hippocampus. These direct recordings have yielded stunning breakthroughs: single neurons within the human orbitofrontal and frontopolar cortices exhibit firing rate modulations that parametrically track the continuous Bayesian variance and counterfactual value of unchosen options in real time, confirming that the latent computational variables extracted from model-based fMRI reflect bona fide cellular-level spiking computations.
Crucially, human intracranial recordings reconcile the long-standing dispute regarding the temporal veracity of model-based fMRI signals. By correlating the high-frequency broadband gamma activity (70–150 Hz)—an exceptional electrophysiological proxy for local multi-unit neuronal firing—with the delayed BOLD responses captured in structural fMRI coordinates, researchers can construct definitive forward models of human neurovascular coupling. This multi-scale alignment confirms that the parametric BOLD modulations observed in Montague’s Leapfrog studies represent genuine energetic expenditures dedicated to Bayesian state updating rather than vascular epiphenomena.
12.2 Naturalistic and Ecologically Valid Decision-Making Environments
A persistent, valid criticism directed at standard cognitive neuroimaging paradigms concerns their artificial, highly abstracted architecture. In a traditional fMRI Leapfrog Task, a supine subject lies rigidly immobilized within a claustrophobic, deafening bore, viewing simplified two-dimensional geometric shapes on an angled mirror and signaling decisions via repetitive thumb presses on a plastic button box. The ecological leap from an organism foraging for survival across high-dimensional ancestral landscapes to a human subject lying in a magnet bore is vast.
To establish true ecological validity, the next generation of computational neuroeconomics is translating the abstract mathematics of the Leapfrog Task into fully immersive, three-dimensional Virtual Reality (VR) environments. Inside a VR foraging landscape, the mathematical mechanics of leapfrog dynamics are translated into spatially distributed resources: an agent must physically navigate an expansive virtual terrain where patches of food or energy regenerate, deplete, and “leapfrog” in utility based on non-stationary hidden Markov rules. By integrating VR headsets with mobile, wearable neuroimaging modalities—predominantly high-density functional Near-Infrared Spectroscopy (fNIRS) and mobile EEG—researchers can track the neurocomputational correlates of exploration in freely moving, physically navigating human agents.
Wearable fNIRS systems, which utilize dual-wavelength near-infrared light arrays to measure changes in oxygenated and deoxygenated hemoglobin across the cerebral cortex, successfully capture prefrontal hemodynamics while participants actively walk, run, and forage. Comparative studies indicate that the computational strategies deployed during immersive, spatial exploration diverge substantially from stationary computer tasks: physical movement introduces natural energetic locomotion costs that biological agents integrate directly into their dynamic programming Bellman equations, grounding Read Montague’s theoretical constructs in genuine ecological biology.
12.3 Synthesis: Montague’s Vision for Precision Computational Psychiatry
More than three decades after the initial convergence of computer science and neurobiology, Read Montague’s grand scientific vision is culminating in the realization of precision computational psychiatry. For over a century, clinical psychiatry has remained stranded in a pre-paradigmatic state, forced to diagnose mental health pathologies based on subjective clinical interviews and descriptive phenomenological manuals (such as the DSM-5) that group biologically heterogeneous diseases under unified, subjective umbrella terms. Two patients diagnosed with “Major Depressive Disorder” may share zero overlapping biological etiologies.
The Leapfrog Task exemplifies how computational neuroeconomics provides the objective mathematical toolkit required to shatter this diagnostic impasse. Under Montague’s vision, a patient presenting with affective, cognitive, or addictive symptoms does not receive a subjective clinical label. Instead, the patient completes a standardized battery of computational assays—with the Leapfrog Task serving as the premier probe of dynamic non-stationary arbitration—while high-resolution multimodal neuroimaging records their underlying neurocircuitry. The behavioral and neural time series are inverted through normative hierarchical models to extract the patient’s precise computational fingerprint: an individualized vector of latent parameters quantifying their learning rates, Bayesian belief precision, counterfactual drift sensitivity, decision conflict thresholds, and frontostriatal connectivity weights.
By mapping an individual’s computational fingerprint against vast, population-level normative reference distributions, clinicians can pinpoint the exact algorithmic failure mode driving the patient’s real-world functional impairment. A patient displaying compulsive drug-seeking driven by a collapse in directed exploration bonuses ($phi$) requires an entirely distinct pharmacological and psychotherapeutic intervention than a patient displaying compulsive behavior driven by an inflated learning rate for negative outcomes ($\alpha^-$). Furthermore, these computational phenotypes serve as prospective prognostic biomarkers, accurately predicting an individual’s therapeutic responsiveness to targeted psychopharmacology, deep brain stimulation (DBS) protocols, or computationally designed cognitive-behavioral therapies. In this synthesis, the Leapfrog Task achieves its ultimate historical potential: transcending laboratory abstraction to establish a foundational pillar in the humane, quantitative, and biologically grounded science of the human mind.
Conclusion
The Leapfrog Task, conceived through the pioneering vision of P. Read Montague, stands as a watershed achievement in the evolution of computational neurobiology and cognitive neuroimaging. By moving decisively beyond the simplistic architectures of static decision theory, the Leapfrog paradigm forces cognitive science to confront the true computational reality of human survival: the imperative to navigate non-stationary, uncertain, and dynamically shifting environments through the delicate arbitration of the explore-exploit trade-off. Through its rigorous mathematical formalization of latent leapfrog hazards, Bayesian belief updates, and epistemic information bonuses, the task provides an unprecedented analytical lens for dissecting the hidden computational machinery of the human brain.
Navigating the extraordinary methodological challenges of task-based fMRI—from temporal deconvolution dilemmas and susceptibility-induced ventral signal dropout to the intricacies of model-based parametric GLMs and dynamic causal modeling—has compelled the neuroimaging community to invent more sophisticated, rigorous, and statistically sound tools. These tools have unveiled a magnificent, distributed neural symphony: the ventromedial prefrontal cortex tracking common currency expected utility; the lateral frontopolar cortex tirelessly holding and monitoring counterfactual alternatives; the dorsal anterior cingulate cortex detecting conflict and signaling volatility-driven policy shifts; the dorsolateral prefrontal cortex maintaining working memory traces and enforcing top-down control; and the striatum and midbrain dopaminergic systems executing prediction error learning and basal ganglia motor gating.
Ultimately, the profound legacy of Montague’s computational neurobiology resonates far beyond the confines of academic neuroimaging laboratories. By transmuting cognitive operations into formal mathematical algorithms, the Leapfrog Task provides a transformative diagnostic engine for computational psychiatry. It offers a future wherein mental illness is unburdened from subjective diagnostic taxonomies and re-established as measurable, treatable variations in the computational parameters of biological wetware. In illuminating how the human brain dares to abandon the security of the known to explore the promise of the unknown, Montague’s work brings us closer to solving one of the most enduring mysteries in science: the physical, biological nature of choice itself.
References
- Behrens, T. E., Woolrich, M. W., Walton, M. E., & Rushworth, M. F. (2007). Learning the value of information in an uncertain world. Nature Neuroscience, 10(9), 1214-1221. https://www.nature.com/articles/nn1888
- Boorman, E. D., Behrens, T. E., Woolrich, M. W., & Rushworth, M. F. (2009). How green is the grass on the other side? Frontopolar cortex and the evidence in favour of alternative courses of action. Neuron, 62(5), 733-743. https://www.cell.com/neuron/fulltext/S0896-6273(09)00392-4
- Daw, N. D., O’Doherty, J. P., Dayan, P., Seymour, B., & Dolan, R. J. (2006). Cortical substrates for exploratory decisions in humans. Nature, 441(7095), 876-879. https://www.nature.com/articles/nature05329
- Eklund, A., Nichols, T. E., & Knutsson, H. (2016). Cluster failure: Why fMRI inferences for spatial extent have inflated false-positive rates. Proceedings of the National Academy of Sciences, 113(28), 7900-7905. https://www.pnas.org/doi/10.1073/pnas.1602413113
- Glover, G. H., Li, T. Q., & Ress, D. (2000). Image-based method for retrospective correction of physiological motion effects in fMRI: RETROICOR. Magnetic Resonance in Medicine, 44(1), 162-167. https://onlinelibrary.wiley.com/doi/abs/10.1002/1522-2594(200007)44:1%3C162::AID-MRM23%3E3.0.CO;2-E
- Kapur, S. (2003). Psychosis as a state of aberrant salience: A framework linking biology, phenomenology, and pharmacology in schizophrenia. American Journal of Psychiatry, 160(1), 13-23. https://ajp.psychiatryonline.org/doi/10.1176/appi.ajp.160.1.13
- McLaren, D. G., Ries, M. L., Xu, G., & Johnson, S. C. (2012). A generalized form of context-dependent psychophysiological interactions (gPPI): A comparison to standard PPI. NeuroImage, 61(4), 1277-1286. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3404477/
- Montague, P. R., Dayan, P., & Sejnowski, T. J. (1996). A framework for mesencephalic dopamine systems based on predictive Hebbian learning. Journal of Neuroscience, 16(5), 1936-1947. https://www.jneurosci.org/content/16/5/1936
- Montague, P. R., Dolan, R. J., Friston, K. J., & Dayan, P. (2012). Computational psychiatry. Trends in Cognitive Sciences, 16(1), 72-80. https://www.nature.com/articles/nn.3036
- Montague, P. R., Hyman, S. E., & Cohen, J. D. (2004). Computational roles for dopamine in behavioural control. Nature, 431(7010), 760-767. https://www.nature.com/articles/nature03015
- O’Doherty, J. P., Hampton, A., & Kim, H. (2007). Model-based fMRI and its application to reward learning and decision making. Annals of the New York Academy of Sciences, 1104(1), 35-53. https://nyaspubs.onlinelibrary.wiley.com/doi/abs/10.1196/annals.1390.002
- Raichle, M. E., MacLeod, A. M., Snyder, A. Z., Powers, W. J., Gusnard, D. A., & Shulman, G. L. (2001). A default mode of brain function. Proceedings of the National Academy of Sciences, 98(2), 676-682. https://www.pnas.org/doi/10.1073/pnas.98.2.676
- Ratcliff, R., & McKoon, G. (2008). The diffusion decision model: Theory and data for two-choice decision tasks. Neural Computation, 20(4), 873-922. https://www.nature.com/articles/nrn2129
- Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. Science, 275(5306), 1593-1599. https://www.science.org/doi/10.1126/science.275.5306.1593
- Shenhav, A., Botvinick, M. M., & Cohen, J. D. (2013). The expected value of control: An integrative theory of anterior cingulate cortex function. Neuron, 79(2), 217-240. https://www.cell.com/neuron/fulltext/S0896-6273(13)00511-4
- Stephan, K. E., Penny, W. D., Daunizeau, J., Moran, R. J., & Friston, K. J. (2009). Bayesian model selection for group studies. NeuroImage, 46(4), 1004-1017. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2799943/