The human brain’s capacity to identify thousands of individual human faces across decades of aging, varied lighting conditions, emotional transformations, and perspective shifts stands as one of the most sophisticated visual feats in the biological world. For decades, cognitive theorists and neurobiologists debated whether this capacity emerges from generalized visual recognition mechanisms tuned to high-dimensional shapes, or whether human visual cortex contains distinct, highly specialized neural modules sculpted by natural selection for the sole purpose of social perception. This theoretical impasse reached a historic turning point in the late 1990s with the advent of functional magnetic resonance imaging (fMRI) and a sequence of rigorous empirical investigations that fundamentally reshaped modern cognitive neuroscience.
In 1997, cognitive neuroscientist Nancy Kanwisher, working alongside colleagues Josh McDermott and Marvin M. Chun, published a groundbreaking paper in The Journal of Neuroscience titled “The Fusiform Face Area: A Module in Human Extrastriate Cortex Specialized for Face Perception.” Utilizing blood-oxygen-level-dependent (BOLD) fMRI, Kanwisher and her team identified an invariant region located in the mid-fusiform gyrus of the human ventral temporal cortex that showed dramatically higher metabolic activity during the visual processing of human faces than during the processing of any other inanimate or animate visual stimulus category. This structure, subsequently designated the Face Fusiform Area (or Fusiform Face Area, commonly abbreviated as the FFA), became the central empirical battleground for debates over cognitive architecture, cortical modularity, perceptual expertise, and the functional organization of the primate brain.
The discovery of the FFA was not merely the identification of a new cortical region; it constituted a methodological and philosophical revolution. Methodologically, Kanwisher pioneered the functional Region of Interest (fROI) localizer approach, an innovative statistical framework designed to circumvent the anatomical blurring caused by classical stereotaxic group averaging. Philosophically, the findings provided the first unambiguous, non-invasive empirical demonstration in neurologically intact humans that the adult neocortex possesses domain-specific processing engines, directly substantiating the theoretical claims of functional modularity. The ensuing decades of research—encompassing intracranial recordings, direct electrical stimulation, high-field neuroimaging, developmental trajectories, non-human primate electrophysiology, and artificial neural networks—have reinforced the seminal nature of Kanwisher’s early experiments while transforming our broader understanding of how the human mind carves reality at its computational joints.
1. Historical Context of Visual Neuroscience and Modularity Prior to 1997
1.1 Early Debates on Cortical Functional Specialization
The conceptual dispute regarding the structural and functional organization of the cerebral cortex represents one of the oldest debates in biological science. In the late nineteenth and early twentieth centuries, the neuroscientific community oscillated violently between two radically opposed paradigms: aggregate field theories (equipotentiality) and extreme cortical localization. The holistic view found its most influential proponent in Karl Lashley, whose extensive lesion studies in rodents led him to formulate the principles of “mass action” and “equipotentiality.” Lashley argued that the rate and efficiency of cognitive performance were functions of the total mass of surviving cortical tissue rather than the destruction of discrete anatomical loci. According to this framework, high-level sensory and cognitive representations were fundamentally distributed throughout widespread associative networks, precluding the possibility of hyper-specialized cortical patches dedicated to singular perceptual categories.
Conversely, the localizationist school traced its lineage back to the clinical observations of Paul Broca and Carl Wernicke in the domain of aphasiology, and to David Ferrier’s precise cortical stimulations in non-human primates. Yet, the question of whether category-specific visual perceptual mechanisms existed remained unresolved. In 1983, the philosopher Jerry Fodor published The Modularity of Mind, which injected rigorous theoretical formalization into this neuroanatomical debate. Fodor hypothesized that the mind comprises computationally autonomous, domain-specific, innately specified, and informationally encapsulated input systems (“modules”) that feed raw perceptual data into unencapsulated, domain-general central processors. A Fodorian module was characterized by mandatory operation, rapid processing speeds, shallow outputs, and dedicated neural architecture. Fodor’s thesis raised an empirical challenge for visual neuroscience: could one identify physical cortical regions that exhibited genuine informational encapsulation and domain specificity for complex visual categories, or was visual processing mediated strictly by an undifferentiated, general-purpose shape engine?
Empirical evidence for face-specialized neural hardware first emerged not in human neuroimaging, but in invasive single-unit electrophysiology in non-human primates. In a series of pioneering experiments conducted during the late 1960s and early 1970s, Charles Gross and his colleagues at Princeton University were recording from neurons within the macaque inferotemporal (IT) cortex. They serendipitously discovered individual units that responded vigorously to complex biological forms, culminating in the historic identification of neurons that fired selectively to depictions of monkey and human faces, but displayed negligible responses to simple geometric forms, bars of light, or complex scrambled patterns. Subsequent studies by Edmund Rolls and David Perrett in the 1980s mapped clusters of face-selective neurons within the banks of the macaque superior temporal sulcus (STS). These animal electrophysiological discoveries provided proof-of-concept that primate evolution had favored dedicated neural circuits for face perception. However, the degree to which the human brain shared this organizational logic remained intensely controversial, as human cognitive anatomy had long diverged down distinct evolutionary and linguistic branches.
1.2 The Emergence of Functional Magnetic Resonance Imaging (fMRI)
Prior to the early 1990s, the empirical investigation of functional specialization in the living human brain was severely restricted by technological bottlenecks. Clinical lesion studies provided invaluable causal insights, but lesions caused by cerebrovascular accidents, trauma, or surgical resections rarely respected functional boundaries, frequently damaging both grey and white matter tracts unpredictably. Early functional neuroimaging technologies, specifically Positron Emission Tomography (PET) and single-photon emission computed tomography (SPECT), offered the first views of human metabolic activity in vivo. However, PET required the intravenous administration of radioactive isotopes (such as H₂¹⁵O), imposing strict radiodosimetric limits that prevented repeated scanning within single human subjects. Furthermore, PET suffered from poor temporal resolution (integrating blood flow over 40 to 90 seconds) and low spatial resolution (often on the order of 8 to 15 millimeters), precluding the fine-grained dissection of small, heterogeneous cortical patches within deep sulcal folds.
The transformation of human cognitive neuroscience began with the discovery of the Blood Oxygenation Level Dependent (BOLD) contrast phenomenon by Seiji Ogawa and colleagues at AT&T Bell Laboratories in 1990. BOLD imaging leverages the differing magnetic properties of oxygenated hemoglobin (diamagnetic) and deoxygenated hemoglobin (paramagnetic). When a localized population of neurons fires, the surrounding vascular architecture responds with a dramatic hyperemic rush of oxygenated blood that exceeds the actual metabolic consumption rate of the active tissue. This influx reduces the relative concentration of deoxygenated hemoglobin, elongates the local transverse relaxation time ($T_2^*$), and amplifies the MR signal intensity captured by gradient-echo echo-planar imaging (EPI) sequences. In 1992, researchers including Ogawa, Kenneth Kwong, and their respective teams published the first successful applications of BOLD fMRI in human subjects, demonstrating robust visual cortex activation in response to simple photic stimulation without radioactive tracers.
Despite this technological leap, the methodological frameworks governing early fMRI experiments in the mid-1990s were poorly equipped to resolve category-specific cortical architectures. Standard analytical workflows relied heavily on spatial smoothing kernels and the averaging of functional data across multiple subjects within standardized stereotaxic coordinate systems, such as the Talairach and Tournoux atlas. These group-averaging approaches systematically blurred distinct, localized functional foci due to the extensive macroanatomical variability of human cortical folding patterns. Early neuroimaging studies investigating face and object perception—most notably those by Justine Sergent in 1992 using PET, and by James Haxby and colleagues in 1994 using fMRI—reported broad patches of activation throughout the ventral occipitotemporal cortex during face viewing. However, these early studies could not decisively adjudicate whether these activations reflected a genuine face-dedicated module, a general-purpose visual categorization network, or the non-specific cognitive effort associated with distinguishing between homogeneous exemplars within an open semantic class.
1.3 Neurological Evidence from Acquired Prosopagnosia
Long before the deployment of magnetic resonance imaging, the most compelling evidence for dedicated face-processing systems came from behavioral neurology. In 1867, Antonio Quaglino published an early report of a patient who, following an apoplectic event, lost the ability to recognize familiar faces while preserving basic intellect and reading capabilities. The condition was formally christened “prosopagnosia” (from the Greek prosopon, meaning face, and agnosia, meaning non-knowledge) by the German neurologist Joachim Bodamer in 1947. Bodamer described two young soldiers who sustained traumatic brain injuries during World War II; despite intact visual acuity, visual fields, and intact object identification, the men were entirely unable to recognize their friends, family members, or even their own reflections in a mirror, relying instead on acoustic cues such as vocal tone or external visual cues such as gait and apparel.
For the next half-century, a fierce clinical debate raged over whether acquired prosopagnosia constituted an autonomous, category-specific deficit or merely a severe manifestation of visual agnosia. Neurologists like Elizabeth Warrington argued that faces simply impose higher perceptual demands than most inanimate objects. Because all human faces share identical first-order configurations—two eyes positioned symmetrically above a central nose, which sits above a horizontal mouth—distinguishing between them requires fine-grained discrimination among subtly varying metric spatial distances and surface properties (second-order relational properties). Inanimate objects, in contrast, frequently differ in their fundamental categorical parts (e.g., a chair has four legs and a flat horizontal back, whereas a teapot has a spout and a handle). Under this “visual expertise” or “perceptual difficulty” view, damage to generic ventral visual processing streams was thought to disproportionately degrade the fine spatial frequency analyses required for face individuation without implicating a face-dedicated neuroanatomical structure.
Conversely, clinical investigators such as Antonio Damasio, Arthur Benton, and Martha Farah accumulated neurobehavioral evidence demonstrating double dissociations between face identification and visual object recognition. Cases were documented in which patients sustained catastrophic visual agnosia for common tools, furniture, and animals, yet retained an unimpaired ability to recognize human faces at the individual level (e.g., patient C.K. studied by Moscovitch, Winocur, and Behrmann). Anatomically, structural computed tomography (CT) and post-mortem neuropathological analyses of prosopagnosic patients revealed that the deficit was most frequently associated with bilateral, or exclusively right-lateralized, lesions of the ventral occipitotemporal cortex, specifically encompassing the lingual and fusiform gyri. Yet, the anatomical boundaries of naturally occurring strokes and surgical resections were notoriously irregular. Neurologists could not precisely localize which microscopic subregion or computational circuit within this expansive territory was directly responsible for the intact generation of face percepts, leaving the definitive anatomical mapping of the human face processor an urgent, unsolved problem in visual cognitive science.
2. Nancy Kanwisher and the Seminal 1997 Investigation
2.1 Conceptual Formulation by Kanwisher, McDermott, and Chun
In the mid-1990s, Nancy Kanwisher, having transitioned from rigorous behavioral investigations of human visual attention and temporal cognition at Harvard and the Massachusetts Institute of Technology (MIT), set out to settle the debate over the functional organization of the ventral visual pathway. Operating at the newly established Massachusetts General Hospital (MGH) NMR Center (now the Athinoula A. Martinos Center for Biomedical Imaging), Kanwisher partnered with two talented researchers: Josh McDermott, then an ambitious graduate student, and Marvin M. Chun, a postdoctoral fellow. The trio conceived an experimental and analytical framework designed to definitively test the hypothesis that human extrastriate cortex contains an area dedicated exclusively to processing faces.
The conceptual logic of their approach rested on an uncompromisingly strict definition of domain specificity. Kanwisher postulated that to validate the modularity hypothesis, one could not merely demonstrate that a cortical region responds strongly to faces; one had to prove that the candidate area responded significantly more strongly to faces than to any other visually complex category, and that this elevated activation could not be explained away by non-specific visual confounders such as visual contrast, luminance, low-level spatial frequencies, or generic attention. Furthermore, they recognized that the conventional fMRI analytical conventions of their day—namely, averaging data across multiple individuals in standardized stereotaxic space—were fundamentally flawed because they washed away functional precision. By treating each individual human brain as an independent architectural entity, Kanwisher and her team established the experimental foundation for what would become their landmark 1997 paper in The Journal of Neuroscience.
The intellectual boldness of the Kanwisher, McDermott, and Chun project lay in its willingness to challenge the prevailing connectionist and domain-general orthodoxy of the computational cognitive science community. Connectionist parallel distributed processing (PDP) models, which dominated cognitive psychology in the 1980s and 1990s, posited that mental representations arise from the continuous, overlapping, and distributed weights of broad neural networks, viewing modular concepts as anachronistic throwbacks to nineteenth-century phrenology. Kanwisher’s enterprise was therefore both empirically risky and philosophically transformative: if a clean, discrete functional region dedicated strictly to a single visual category could be unmasked using cutting-edge neuroimaging, the foundational architecture of the human mind would have to be understood through the lens of specialized computational components.
2.2 The Functional Region of Interest (fROI) Methodological Revolution
The single most consequential methodological innovation introduced by Kanwisher and her collaborators was the formulation and rigorous application of the functional Region of Interest (fROI) method. Standard neuroimaging analyses historically aligned all experimental subjects’ structural scans into a common stereotaxic space, such as the Talairach coordinate system or the Montreal Neurological Institute (MNI) template. In this classical framework, statistical parametric maps (SPMs) were generated by averaging functional voxel values across the group. However, human primary and associative cortices exhibit extraordinary macroanatomical and microstructural heterogeneity: sulcal and gyral trajectories vary widely between individuals, meaning that a fixed stereotaxic coordinate (e.g., $x=40, y=-55, z=-10$) might correspond to the lateral fusiform gyrus in one participant, the collateral sulcus in another, and the inferior temporal sulcus in a third.
Kanwisher argued that this group-averaging methodology inflicted catastrophic damage on functional data: it smeared discrete, localized functional patches over vast swaths of cortical space, artificially diminished statistical power, and obscured subtle functional boundaries between adjacent, computationally divergent regions. To eliminate this issue, Kanwisher instituted an independent functional localizer protocol. Each subject underwent an independent scanning run—the “localizer run”—designed to contrast the metabolic response during face viewing with that during common object viewing. This localizer was used strictly to pinpoint the precise anatomical voxels exhibiting face-selective responses within that specific subject’s individual cortical landscape, accounting for the unique anatomical morphology of their ventral temporal lobe.
Once an fROI was statistically defined within an individual participant using this independent run, its spatial coordinates and cluster boundaries were locked. All subsequent experimental hypotheses—testing responsiveness to scrambled faces, inverted faces, body parts, animals, and low-level visual controls—were evaluated strictly on independent datasets collected in separate scanning runs from the same subject. By extracting parameter estimates (percent BOLD signal change) solely from these a priori defined fROIs, Kanwisher’s method achieved three profound advantages:
- It avoided circular statistical reasoning (the so-called “double-dipping” fallacy later formalized in modern neuroimaging literature).
- It maximized functional sensitivity by aggregating signal only from authentically face-selective neural tissue.
- It preserved the true functional response profile of the human brain at single-subject resolution without relying on the assumption of identical macroanatomical mapping across the human species.
2.3 Core Objectives and Hypotheses Tested in the 1997 Study
The primary empirical objective of the 1997 study was to systematically evaluate and eliminate competing alternative hypotheses regarding the functional profile of the ventral temporal activation elicited by faces. Kanwisher, McDermott, and Chun sought to establish whether the target cortical patch fulfilled four demanding operational criteria:
- Category Specificity: Does the candidate region exhibit significantly greater metabolic activation for faces than for inanimate, familiar, three-dimensional objects matched for visual variety?
- Invariance to Low-Level Visual Features: Is the activation driven by primary physical properties of the image—such as spatial frequency spectra, visual contrast, edge density, or luminance—or by the higher-order geometric configuration that defines a face?
- Holistic vs. Feature-Based Processing: Does the localized region respond to individual, isolated facial components, or does it require the canonical, upright spatial relationship among features (the relational face template) to achieve maximal engagement?
- Replicability and Stability: Can the localized response be consistently detected within the same individuals across distinct experimental runs, and can it be replicated across separate cohorts of participants scanned on different days?
By designing a succession of nested control conditions—systematically comparing intact faces against two-dimensional line drawings, phase-scrambled faces, inverted faces, and biological control stimuli—the researchers aimed to establish an unambiguous double dissociation between face recognition and general visual object identification. They hypothesized that if a dedicated module existed, it would occupy a consistent topological niche along the ventral processing stream, demonstrate selective vulnerability to structural manipulations that abolish the behavioral “face inversion effect,” and refuse to respond with equal magnitude to any control category, no matter how visually complex or structurally demanding.
3. Experimental Paradigms and Methodological Rigor in the Landmark Experiments
3.1 Stimulus Selection and Controlled Visual Sets
Achieving methodological validity in high-level visual neuroscience requires meticulous control over visual properties that could inadvertently induce differential visual cortex activation. In their landmark 1997 series of experiments, Kanwisher and colleagues curated stimulus sets designed to rule out low-level confounds. The primary experimental category consisted of digitized, high-contrast, grayscale photographs of human faces displaying neutral emotional expressions. These faces featured both male and female identities, cropped to eliminate peripheral non-facial cues such as distinct hairstyles, jewelry, and clothing collar lines, ensuring that participants were processing facial features rather than idiosyncratic silhouettes.
The critical comparison category consisted of common, familiar inanimate objects. These objects—ranging from houses, chairs, and musical instruments to tools, appliances, and motor vehicles—were matched to the face stimuli in fundamental two-dimensional image characteristics. Luminance distributions were normalized across sets to prevent global illumination disparities from driving early visual retinotopic responses. Contrast metrics (specifically root-mean-square contrast) were calculated and standardized. Furthermore, to evaluate whether the region responded to complex shapes rather than face-specific architectures, the team constructed a control stimulus class of “scrambled faces.” These were created by cutting face images into a fine mosaic grid and randomly permuting the spatial positions of the tiles, or by scrambling the phase spectra of the two-dimensional Fourier transforms while preserving the original amplitude spectra.
These scrambled face stimuli possessed the identical spatial frequency distribution, global luminance, and total contrast energy of intact faces, yet their semantic meaning and structural configuration were utterly destroyed. If the ventral temporal activations documented in early imaging studies were merely artifacts of the unique spatial frequency content of face photographs (which feature pronounced power along specific horizontal and vertical orientations), the scrambled faces should have evoked a BOLD response identical to that of intact faces. If, however, the area was tuned to high-level visual configurations, the response to scrambled faces would collapse to baseline levels.
3.2 Block Design Architecture and Scanning Protocols
The temporal architecture of the fMRI experiments employed a standard block design (or boxcar design), which maximized statistical power and blood-oxygen-level-dependent contrast-to-noise ratios. Experiments were organized into continuous functional runs, typically lasting between five and six minutes. Each run comprised alternating experimental epochs (blocks) lasting 30 seconds each, separated by interleaved baseline blocks of fixational rest during which participants stared at a central fixation point on an otherwise blank screen.
Within each 30-second stimulus block, approximately 45 individual visual stimuli belonging to a single category (e.g., intact faces, common objects, scrambled faces, or inverted faces) were presented in rapid succession. Stimuli were flashed for brief durations—typically 500 milliseconds—followed by an inter-stimulus interval of roughly 250 milliseconds. This rapid serial visual presentation ensured that participants maintained engagement while preventing saccadic eye movements from systematically scanning disparate quadrants of the display. The order of the stimulus blocks was counterbalanced across runs (e.g., ABACAB vs. ACABAC) to eradicate potential confounding effects driven by scanner signal drift, vascular inertia, or neural adaptation across time.
Imaging was executed on a 1.5 Tesla General Electric Signa clinical MRI scanner retrofitted with high-performance echo-planar imaging gradients. Structural localization was acquired using high-resolution $T_1$-weighted anatomical pulse sequences (such as a 3D SPGR sequence), providing thin, sub-millimeter axial or sagittal partitions of the whole brain. Functional BOLD acquisitions utilized single-shot gradient-echo echo-planar imaging pulse sequences with a repetition time (TR) ranging between 2.0 and 3.0 seconds, an echo time (TE) of approximately 30 to 40 milliseconds, a flip angle of 90 degrees, and slice thicknesses spanning 4 to 6 millimeters. The functional imaging volume was tilted to orient parallel or oblique to the ventral temporal cortical surface, ensuring optimal anatomical coverage of the occipitotemporal junction and the fusiform gyrus.
3.3 Behavioral Tasks and Attentional Modulation During fMRI
A classic vulnerability of early neuroimaging experiments was the reliance on passive viewing paradigms. When human participants passively view visual displays, unconstrained cognitive states can distort empirical findings: faces might naturally evoke heightened intrinsic interest, capturing spontaneous attention and emotional appraisal to a far greater extent than inanimate items such as tables or wrenches. Under passive viewing, an investigator cannot easily determine whether elevated BOLD responses reflect domain-specific category selectivity or merely asymmetric attentional allocation and heightened arousal.
To definitively untangle attentional modulation from category-specific sensory processing, Kanwisher, McDermott, and Chun implemented active behavioral monitoring tasks directly within the magnet. The most important of these was the “one-back” image repetition-detection task. During one-back functional blocks, participants were instructed to monitor the rapid visual stream and press a pneumatic or fiber-optic response button whenever the exact same image appeared consecutively twice in a row. These immediate repetitions occurred randomly 2 to 4 times per 30-second block, forcing participants to sustain an identical, vigilant focus across all conditions, whether they were viewing faces, scrambled shapes, houses, or chairs.
Behavioral performance metrics, specifically reaction times and discriminability indices ($d’$), were recorded and statistically matched across categories. The results were conclusive: even when performance accuracy was completely balanced across conditions, and even when an identical one-back behavioral rule was applied across every stimulus class, the fusiform region continued to exhibit its robust face-selective activation. Furthermore, subsequent control paradigms demonstrated that passive viewing, one-back tasks, and visual categorization tasks yielded nearly identical response profiles within the region. This confirmed that the observed signal difference was fundamentally driven by the physical and conceptual input properties of the visual stimuli, rather than an artifact of voluntary attentional focus or cognitive effort.
4. Functional Localization and Characterization of the FFA
4.1 The Subtraction Method: Faces Versus Common Objects
The foundational proof of face-selective cortical specialization was established through the classic cognitive subtraction method applied to functional magnetic resonance data. By contrasting the hemodynamic response elicited during the processing of intact human faces against the response elicited during the processing of familiar common objects (contrast: $[Faces – Objects]$), Kanwisher and her colleagues isolated neural regions whose metabolic expenditure was significantly heightened for faces after subtracting out the metabolic processes common to general shape processing, spatial attention, and semantic object categorization.
The statistical subtraction revealed an exceptionally clear, consistent functional locus situated along the lateral aspect of the fusiform gyrus. In the seminal 1997 study, this face-selective activation cluster was successfully identified in 12 out of the 15 scanned participants. The activation was characterized by extraordinarily high statistical significance, with individual voxel clusters surpassing conservative statistical thresholds (typically $p < 10^{-4}$ to $p < 10^{-6}$, uncorrected, within individual brains). The magnitude of the response divergence was immense: parameter estimates derived from the independent functional runs demonstrated that the lateral fusiform region responded roughly twice as intensely to faces as it did to common three-dimensional objects.
This localized cortical patch, formally designated the Face Fusiform Area (FFA), was not an amorphous, widespread territory, but a relatively circumscribed functional cluster. In most participants, the FFA spanned an average volume of several hundred cubic millimeters (typically 1 to 2 cubic centimeters of cortical grey matter, depending on statistical thresholding and individual morphology). The specificity of this response was remarkable: while wide territories across the lateral occipital complex responded robustly to inanimate objects, the FFA displayed a sharp, step-like increase in metabolic recruitment exclusive to the face category, providing the first localized neuroimaging evidence for domain-specific visual modularity in the human brain.
4.2 Controlling for Low-Level Visual Features: Scrambled and Inverted Faces
To eliminate the possibility that the FFA was simply responding to low-level spatial frequency profiles or local edge contrasts, Kanwisher, McDermott, and Chun evaluated the region’s response to phase-scrambled and mosaic-scrambled faces. In every individual subject who possessed a robustly localized FFA, the BOLD response to scrambled faces was entirely extinguished, plunging down to the level of fixational rest baselines. Because scrambled faces preserved the exact two-dimensional Fourier energy spectra, total luminance, and contrast variations of intact faces, this collapse confirmed that the FFA is fundamentally blind to disorganized low-level physical properties, requiring a recognizable geometric visual structure to fire.
The researchers subsequently tested one of the most famous behavioral phenomena in perceptual psychology: the Face Inversion Effect, originally documented by Robert Yin in 1969. Yin observed that turning an image upside down severely impairs human recognition memory for faces, whereas the recognition of other visually complex objects (such as houses, airplanes, or landscapes) suffers only minor decrements. This asymmetric behavioral deficit is widely interpreted as evidence that upright faces are perceived “holistically”—that is, by integrating the relational spacing among features into a single, unified perceptual gestalt—whereas inverted faces must be parsed piece-by-piece using piecemeal, feature-based analytical routines.
When Kanwisher and colleagues presented inverted faces within their independent functional localizer paradigms, the fROI showed a significant, measurable reduction in BOLD amplitude compared to upright faces. While inverted faces still elicited higher activation than common inanimate objects (suggesting that the upside-down stimuli still partially engaged the face template), the response was substantially attenuated relative to canonical upright presentations. This marked attenuation within the FFA mirrored the classic behavioral inversion decrement, demonstrating that the FFA is specifically tuned to the upright, holistic spatial arrangement of human facial anatomy, and solidifying its identity as the functional substrate mediating holistic perceptual expertise.
4.3 Controlling for Visual Complexity and Human Body Elements
An alternative critique of the modular hypothesis proposed that the FFA was not a face processor per se, but an animate-category processor, or a detector of human biological forms in general. To test this hypothesis, the experimental protocol was expanded to include visual controls consisting of non-face human body parts. Kanwisher scanned participants viewing photographs of human hands cropped and presented under identical lighting, contrast, and scaling parameters as the face stimuli.
The results provided another critical dissociation. Human hands failed to evoke significant metabolic responses within the FFA: the BOLD signal elicited by hands was statistically indistinguishable from, or only slightly higher than, the signal elicited by inanimate household objects, while remaining dramatically below the response elicited by faces. This proved that the FFA is not an undifferentiated “human body area” or a general biological agent detector. (Indeed, several years later in 2001, Kanwisher and Paul Downing discovered an entirely distinct, adjacent region in the lateral occipitotemporal cortex dedicated to human bodies: the Extrastriate Body Area, or EBA).
The investigators also tested biological specificity by presenting front-facing animal faces (such as the faces of cats and dogs). While animal faces evoked moderately elevated responses within the human FFA relative to inanimate objects—indicating that the region’s internal template generalizes to mammalian facial structures displaying bilateral symmetry, dual eyes, and a snout—the magnitude of activation was systematically lower than that evoked by conspecific human faces. Finally, to eliminate the alternative explanation that the FFA is simply driven by visually complex items possessing numerous intricate component parts, the authors contrasted faces against images of complex machinery and architectural elevations. These visually intricate inanimate structures uniformly failed to activate the FFA. Through this exhaustive process of elimination, Kanwisher and her colleagues demonstrated that the FFA fulfilled the operational criteria of an informationally encapsulated, domain-specific visual module.
5. Anatomical Properties, Lateralization, and Structural Topography
5.1 Anatomical Landmarks of the Lateral Fusiform Gyrus
The human fusiform gyrus (also historically designated the lateral occipitotemporal gyrus) is an expansive, elongated anatomical structure spanning the ventral aspect of the temporal and occipital lobes. It is bordered medially by the deep collateral sulcus (CoS), which separates the fusiform gyrus from the parahippocampal gyrus and the lingual gyrus, and laterally by the occipitotemporal sulcus (OTS), which demarcates the boundary between the fusiform gyrus and the inferior temporal gyrus. The macroanatomical terrain of this ventral surface is notorious for its extensive individual morphological variability.
Within this broader territory, the FFA is predominantly anchored to the mid-fusiform gyrus. In recent years, high-resolution anatomical and functional profiling pioneered by Kevin Weiner and Kalanit Grill-Spector has identified a shallow, longitudinal sulcus that bisects the fusiform gyrus along its anterior-posterior axis: the mid-fusiform sulcus (MFS). The MFS serves as a reliable macroanatomical landmark for the functional localization of face selectivity. Specifically, the face-selective cortex is systematically organized in relation to the anterior tip and fundus of the MFS, providing an anatomical anchor across human subjects despite surrounding gyral variability.
Furthermore, detailed spatial mapping has demonstrated that what was originally characterized as a single monolithic entity (the FFA) is frequently resolvable into two distinct, anatomically adjacent functional patches along the anterior-posterior axis:
- FFA-1: The posterior patch, often situated along the posterior aspect of the mid-fusiform sulcus.
- FFA-2: The anterior patch, positioned more rostrally, typically near the anterior terminus of the MFS.
These two patches exhibit subtle differences in their receptive field architectures, anatomical connectivity, and downstream computational roles, with FFA-2 demonstrating greater invariance to viewpoint transformations and stronger functional coupling to anterior temporal language and identity networks.
5.2 Hemispheric Asymmetry and Right-Hemisphere Dominance
A consistent neuroanatomical characteristic of the FFA is its marked hemispheric asymmetry. In the vast majority of human participants, the FFA demonstrates clear right-hemisphere dominance. When functional localizers are applied, the right FFA (often designated rFFA) is nearly ubiquitous, identifiable in roughly 90% to 98% of healthy right-handed individuals. It typically occupies a substantially larger cortical volume, exhibits significantly higher BOLD percent signal changes, and withstands much more stringent statistical thresholding than its contralateral counterpart.
The left FFA (lFFA), while frequently present, is identifiable in a smaller proportion of subjects (historically between 60% and 85% of scanned cohorts). It tends to be smaller in spatial extent and displays a less pronounced functional contrast when faces are subtracted against non-face objects. Neuropsychological lesion literature matches this asymmetry: unilateral damage to the right ventral occipitotemporal cortex is frequently sufficient to induce profound, chronic acquired prosopagnosia, whereas unilateral damage to the left ventral occipitotemporal cortex more commonly results in pure alexia (the inability to read words despite preserved general intellect) or mild visual agnosias, sparing core face recognition capabilities.
Cognitive and evolutionary neuroscientists have proposed several complementary hypotheses to account for this right-hemispheric dominance. One classic formulation posits an evolutionary division of labor within the ventral stream: the left hemisphere, which in most humans is dedicated to linguistic processing, speech production, and the sequential syntactic parsing of discrete symbols, biased the left ventral occipitotemporal cortex toward the fine-grained, analytic parsing of orthography (leading directly to the emergence of the Visual Word Form Area, or VWFA). Conversely, the right hemisphere, relatively liberated from communicative orthographic computational constraints, maintained a bias toward holistic, gestalt, parallel visuospatial configurations, optimizing its neural tissue for the invariant geometric demands of face and scene processing.
5.3 Inter-Subject Variability in FFA Morphology
One of the primary reasons the FFA eluded early PET and low-resolution fMRI investigations was the profound inter-subject variability in its precise morphological instantiation. While the FFA is consistently localized within the mid-fusiform territory, its precise coordinates within standardized stereotaxic space (such as Talairach or MNI) show marked dispersion across individuals. In standard MNI space, the coordinates for the right FFA typically cluster around $x = 38$ to $44$, $y = -48$ to $-60$, and $z = -12$ to $-20$. However, individual centroids can easily deviate by 10 to 15 millimeters along any cardinal axis—a distance that, in the brain, spans entirely distinct cytoarchitectonic areas.
This anatomical dispersion arises from several sources:
- Variable Gyral Folding: The mid-fusiform sulcus can be continuous, bifurcated, fragmented into dual segments, or absent altogether in a subset of healthy individuals.
- Spatial Extent Differences: The functional volume of the FFA varies dramatically across healthy cohorts, ranging from under 200 cubic millimeters to well over 1,500 cubic millimeters, reflecting differing degrees of cortical surface allocation.
- Depth and Cortical Curvature: The active cluster can sit squarely on the lateral crest of the fusiform gyrus, wrap into the lateral bank, or descend into the depths of the adjacent occipitotemporal sulcus.
Because of this spatial variance, any statistical approach that relies on averaging unaligned functional data across twelve or twenty participants in standardized space inevitably averages face-selective voxels with non-selective voxels from adjacent tissues (such as general object-responsive lateral occipital tissue or medial scene-selective parahippocampal tissue). The resulting group-averaged activation map yields a smeared, low-amplitude patch of ambiguous significance. Nancy Kanwisher’s rigorous insistence on individual-subject functional region-of-interest analysis was therefore essential to bypass this morphological noise, demonstrating that beneath the superficial heterogeneity of cortical folding lies a highly conserved, invariant functional architecture.
6. The Expertise Hypothesis vs. Domain-Specificity Debate
6.1 Gauthier’s Perceptual Expertise Model and the ‘Greeble’ Studies
The modular, domain-specific interpretation championed by Kanwisher met fierce theoretical and empirical resistance in the late 1990s and early 2000s. The most prominent opposition was mounted by Isabel Gauthier, Michael Tarr, and their collaborators, who formulated the “Perceptual Expertise Hypothesis.” Gauthier contended that the FFA was not an innately specified module dedicated exclusively to the social category of faces; instead, she argued that the region constitutes a general-purpose cortical processor optimized for fine-grained, subordinate-level visual discrimination among highly homogeneous visual exemplars—an engine for visual expertise.
According to the expertise model, humans are not born with a face module; rather, through thousands of hours of social experience accumulated from infancy onward, faces become the ultimate visual category for which every neurologically intact human acquires deep perceptual expertise. Because all faces share the same basic first-order features, recognizing an individual face requires discriminating between subtle variations in second-order metric relationships (e.g., the precise distance between the pupil and the bridge of the nose). In contrast, common visual objects are typically identified at the basic semantic level (e.g., “car,” “dog,” “table”), which requires only crude part-based parsing. Gauthier asserted that if an adult human were to acquire equivalent subordinate-level expertise for any non-face visual category, the FFA would become equally recruited for that category.
To empirically test this hypothesis, Gauthier and colleagues designed artificial, computer-generated three-dimensional visual entities christened “Greebles.” Greebles were novel, symmetrical forms that possessed a common spatial configuration: each Greeble featured a vertically oriented central body and four distinctive appendages (termed “boges,” “quargs,” and “krups”) arranged in invariant positions. The researchers trained human participants over extensive sessions lasting several weeks to categorize Greebles at the subordinate level (identifying individual Greeble “identities” and “genders”). Gauthier and colleagues reported that as participants attained behavioral criteria for expertise, their BOLD signal within the anatomically localized FFA significantly increased for Greebles compared to naive controls, with Greebles also eliciting a behavioral inversion effect. These findings were promoted as decisive proof that the FFA is fundamentally an expertise engine rather than a domain-specific face module.
6.2 Kanwisher’s Rebuttals and Direct Counter-Experiments
Nancy Kanwisher launched an immediate and comprehensive empirical counter-offensive against the expertise hypothesis, pointing out critical methodological and conceptual flaws in the Greeble studies. First, Kanwisher noted that the Greeble experiments frequently failed to employ strict, independent fROI localizer protocols: the reported activations were often identified via group-averaging methods or lenient functional contrasts that blended face-selective voxels with neighboring object-responsive cortical tissue. Second, Kanwisher emphasized that Greebles looked strikingly biological, possessing symmetrical, upright appendages arranged in a layout that strongly resembled eyes, a nose, and a mouth. As a consequence, it was plausible that Greebles were simply parasitizing the pre-existing, face-selective neural template because of their resemblance to faces, rather than recruiting the area via generic visual expertise.
To adjudicate this controversy, Kanwisher, Kalanit Grill-Spector, and their collaborators carried out a series of definitive studies examining real-world, natural visual experts. If the expertise hypothesis were correct, individuals who possessed years of intensive subordinate-level categorization expertise with real-world non-face objects—such as master car mechanics or passionate ornithologists (bird watchers)—should exhibit massive, face-equivalent FFA recruitment when viewing cars or birds, respectively.
In meticulously controlled fMRI experiments published by Kanwisher and colleagues, bird experts and car experts were scanned while viewing faces, birds, cars, and familiar control objects. When the FFA was independently localized using standard, conservative criteria at the individual-subject level, the results delivered a decisive blow to the expertise model:
- Bird experts viewing birds, and car experts viewing cars, showed negligible increases in FFA activation over baseline objects.
- The absolute magnitude of this expertise-related increment was dwarfed by the massive BOLD response elicited by faces in those exact same individuals.
- Subordinate-level categorization of non-face objects recruited lateral occipital and ventral temporal networks outside the FFA, leaving the core FFA almost entirely unresponsive to non-face visual expertise.
Subsequent psychophysical and functional neuroimaging studies by researchers such as Yaoda Xu reinforced Kanwisher’s stance, demonstrating that the behavioral hallmarks of face perception (such as the composite face effect and the holistic inversion effect) did not reliably replicate with Greebles or real-world expertise objects once confounds were removed.
6.3 Theoretical Synthesis of the Modularity vs. Process Debate
The fierce debate between Kanwisher’s domain-specific modular model and Gauthier’s flexible-process expertise framework represents one of the most thoroughly scrutinized intellectual exchanges in contemporary neuroscience. Over two decades of subsequent empirical research have yielded a widespread scientific consensus that firmly favors Kanwisher’s original formulation, albeit with nuanced evolutionary and computational refinements.
The contemporary synthesis distinguishes between two separate questions: the nature of the representation (what the region represents) and the process performed on that representation (how the region computes). Kanwisher successfully demonstrated that the FFA is unequivocally domain-specific in its representational profile: its computational operations are triggered primarily, if not exclusively, by faces. The minuscule modulations in signal occasionally observed for expertise categories in certain studies are widely understood as reflecting top-down attentional amplification or the incidental engagement of adjacent, non-face visual channels, rather than the core functional mandate of the FFA.
Today, cognitive neuroscience acknowledges that while the human brain possesses remarkable neuroplasticity, this plasticity does not imply functional equipotentiality. The FFA does not drift arbitrarily from category to category across the lifespan. Instead, it serves as a dedicated, specialized computational node whose architecture is optimized—both phylogenetically through evolutionary selection and ontogenetically through early developmental visual exposure—to process the unique, highly demanding multidimensional space of conspecific facial identity. The expertise debate ultimately served to harden the scientific community’s methodological standards, solidifying the FFA as the textbook exemplar of cortical functional modularity in the human brain.
7. The Surrounding Ventral Visual Pathway Network
7.1 The Occipital Face Area (OFA) and the Distributed Face Network
Although the Face Fusiform Area rapidly became the centerpiece of human face-processing research, Nancy Kanwisher repeatedly cautioned that the FFA does not operate as an isolated, solitary island within the neocortex. Rather, it represents the primary hub of a distributed, multi-component neural network spanning the ventral and lateral visual processing streams. The first major node positioned upstream from the FFA is the Occipital Face Area (OFA), located in the lateral inferior occipital gyrus (IOG).
Identified via the same functional localizer logic ($[Faces – Objects]$), the OFA exhibits robust face selectivity, but displays distinct functional and computational properties. In classic hierarchical feedforward models—such as the influential distributed framework formulated by James Haxby, M. Ida Gobbini, and Matthew Hoffman in 2000—the OFA is conceptualized as an early structural gatekeeper. It processes the low-level, local visual primitives of a face (individual eyes, nose, mouth) in a retinotopically constrained manner, and subsequently projects these feedforward representations along the ventral temporal pathway to the FFA, where they are synthesized into an integrated, holistic, and identity-preserving gestalt representation.
A third essential component of this core face network resides in the posterior banks of the Superior Temporal Sulcus (pSTS). The Haxby model establishes a fundamental functional dissociation between the ventral stream nodes (OFA and FFA) and the lateral stream node (STS):
- FFA & OFA: Specialize in the invariant, structural aspects of faces—the static anatomical features necessary to establish individual identity across varying conditions.
- Superior Temporal Sulcus (STS): Specializes in the dynamic, changeable aspects of faces—processing gaze direction, lip movements, and transient emotional facial expressions essential for immediate social communication and theory of mind.
7.2 The Parahippocampal Place Area (PPA) Discovery
Shortly after the discovery of the FFA, Nancy Kanwisher, collaborating with her graduate student Russell Epstein in 1998, set out to determine whether the modular architecture of the ventral stream extended to other ecological visual categories. Applying their functional localizer methodology to visual scenes, Epstein and Kanwisher discovered an adjacent, highly selective functional patch situated along the collateral sulcus and parahippocampal gyrus: the Parahippocampal Place Area (PPA).
The PPA demonstrated an absolute category selectivity that formed an exquisite, symmetrical double dissociation with the FFA:
- While the FFA fired vigorously to human faces and fell silent during the presentation of scenes, landscapes, and architectural spaces, the PPA did the exact opposite.
- The PPA responded maximally to visual depictions of real-world environmental scenes—such as photographs of city streets, natural landscapes, and empty furnished rooms—while showing dramatic metabolic drops during the presentation of faces, objects, or isolated elements.
The discovery of the PPA was a monumental confirmation of Kanwisher’s modular vision. It proved beyond doubt that the FFA was not an eccentric neuroanatomical anomaly, but part of a systematic organizational logic governing the human ventral visual pathway. Just as social navigation required a dedicated computational engine for conspecific identification (the FFA), spatial navigation required a dedicated computational engine for extracting topological layout, boundary geometry, and environmental affordances (the PPA). The adjacent coexistence of these two functionally pure, category-selective regions within the ventral temporal lobe fundamentally transformed our understanding of extrastriate visual architecture.
7.3 The Extrastriate Body Area (EBA) and Lateral Occipital Complex (LOC)
In subsequent years, the empirical exploration of the visual ventral stream revealed a broader topography of specialized and general-purpose visual processors. In 2001, Paul Downing, Yuhong Jiang, Miles Shuman, and Nancy Kanwisher published the discovery of the Extrastriate Body Area (EBA). Located in the posterior inferior temporal sulcus/middle temporal gyrus, the EBA responded selectively to human bodies and isolated body parts (limbs, torsos, hands) relative to faces, objects, and animals. Later, a second body-selective region—the Fusiform Body Area (FBA)—was identified directly adjacent to, and partially overlapping with, the lateral fusiform face area, establishing parallel processing pathways for static bodily forms and social postures.
These category-selective outposts were situated within an overarching visual architecture that also contained large-scale, general-purpose shape recognition structures. Most notable among these is the Lateral Occipital Complex (LOC), originally delineated by Rafael Malach and colleagues in 1995. The LOC spans the lateral aspect of the occipital lobe extending into the ventral stream, exhibiting elevated BOLD activation whenever participants view structured, three-dimensional objects (tools, geometrical solids, artifacts) as opposed to texture patterns or random visual noise.
The structural layout of these regions revealed fundamental topographic principles along the ventral temporal cortex. As demonstrated by Rafael Malach and his team (including Uri Hasson and Ifat Levy), category selectivity maps onto an underlying eccentricity gradient:
- Center-Biased Processing: Visual categories that require high-acuity foveal inspection—such as faces (the FFA) and words (the VWFA)—are localized along the lateral fusiform gyrus, which is retinotopically linked to central foveal representations.
- Periphery-Biased Processing: Visual categories that encompass expansive visual fields and peripheral vision—such as spatial scenes and architectural layouts (the PPA)—are systematically localized along the medial aspect of the ventral cortex, linked to peripheral visual space.
8. Converging Methodological Evidence: Lesions, Electrophysiology, and Direct Stimulation
8.1 Intracranial Electroencephalography (iEEG) and Electrocorticography (ECoG)
While functional magnetic resonance imaging provided unprecedented spatial localization in healthy human subjects, it remained susceptible to two fundamental criticisms: it lacked sub-second temporal resolution, and, as a correlational imaging technique, it could not definitively establish causal necessity. If metabolic activation in the FFA was an epiphenomenal byproduct of downstream associative processing, fMRI alone could not prove that the FFA was computationally necessary for the conscious perception of a face.
Crucial converging evidence emerged from invasive clinical neurophysiology. Patients with pharmacoresistant epilepsy frequently undergo surgical implantation of subdural electrocorticographic (ECoG) grids or depth stereotactic electroencephalography (sEEG) electrodes along the ventral temporal cortex to locate the epileptogenic zone. In a succession of landmark recordings conducted by Truett Allison, Gregory McCarthy, Aina Puce, and later by Jacques Jonas and Bruno Rossion, intracranial electrodes positioned directly on the human fusiform gyrus recorded localized local field potentials during visual tasks.
These intracranial recordings uncovered a massive, sharply tuned negative potential occurring precisely 200 milliseconds following stimulus onset—the intracranial N200 component. The intracranial N200 demonstrated category selectivity that aligned with the fMRI-defined FFA: it fired vigorously to faces, but showed virtually no deflection for common inanimate objects, scrambled images, or non-face body parts. Furthermore, these recordings revealed the millisecond-by-millisecond temporal unfolding of face perception, demonstrating that face-selective signals in the human fusiform gyrus emerge rapidly and autonomously between 150 and 200 milliseconds post-stimulus, eliminating the concern that FFA activation was merely a late top-down artifact of visual imagery or post-perceptual semantic memory.
8.2 Direct Electrical Cortical Stimulation (ECS) in Conscious Patients
The definitive causal proof of the FFA’s direct necessity for conscious face perception came from direct electrical cortical stimulation (ECS) performed in awake human patients. In 2012, Josef Parvizi and his research group at Stanford University—collaborating with Nancy Kanwisher and Kalanit Grill-Spector—conducted an experiment on an epileptic patient implanted with subdural electrodes over the exact anatomical territory of the right FFA.
The researchers first utilized high-resolution BOLD fMRI to functionally localize the patient’s right FFA, confirming that the intracranial ECoG electrode contacts rested directly atop the face-selective cluster along the lateral fusiform gyrus. While the conscious patient was looking at the face of the neurologist, the researchers delivered a mild electrical current through the pair of electrodes overlying the FFA without warning the patient of the exact delivery timing. The patient immediately gasped and reported a profound, immediate visual illusion: the neurologist’s face seemed to transform, warp, and melt into that of a different, unrecognizable individual.
The patient memorably described the phenomenon:
“You just turned into somebody else. Your face metamorphosed. Your nose got pushed to the left, and you looked like somebody else… only your face changed, everything else stayed the same.”
Critically, when the researchers stimulated adjacent, non-face-selective electrodes on the fusiform gyrus, no facial distortions occurred. Furthermore, when the FFA was electrically stimulated while the patient looked at non-face visual objects (such as an inanimate object or an abstract line pattern), the patient reported no visual distortion of the object; the metamorphopsia was triggered exclusively when looking at human faces. This experiment provided causal verification of Kanwisher’s modular hypothesis: the FFA is not an epiphenomenal spectator, but a functionally necessary, causally operative computational node that generates conscious perceptual representations of human faces.
8.3 Transcranial Magnetic Stimulation (TMS) Over Face-Selective Regions
To establish non-invasive causal evidence in healthy cohorts, cognitive neuroscientists turned to Transcranial Magnetic Stimulation (TMS). Because the FFA is situated deep within the ventral temporal cortex along the base of the skull, delivering focused magnetic pulses directly to the mid-fusiform gyrus via external surface coils is technically unfeasible without stimulating intervening temporal musculature and peripheral nerves. However, the upstream node of the network—the Occipital Face Area (OFA)—is situated near the lateral cortical surface in the inferior occipital gyrus, making it accessible to transcranial magnetic intervention.
In a series of studies led by David Pitcher, Vincent Walsh, and Bradley Duchaine, repetitive and event-related double-pulse TMS was applied over the right OFA while participants performed visual discrimination tasks involving faces, bodies, and inanimate objects. The delivery of TMS specifically impaired performance on face-matching tasks, leaving body-part discrimination and object-matching performance intact. In contrast, delivering TMS over the adjacent Extrastriate Body Area (EBA) selectively degraded body discrimination while leaving face perception unperturbed.
Furthermore, chronometric TMS experiments—delivering magnetic pulses at precise time intervals following image onset—demonstrated that disrupting the OFA at early latency windows (between 60 and 100 milliseconds post-stimulus) caused the most catastrophic impairments in downstream face discrimination. These findings validated the hierarchical nature of the network, showing that the intact computational output of early lateral visual nodes is sequentially routed to the ventral fusiform cortex to support holistic identity recognition, providing causal triangulation across clinical lesions, invasive stimulation, and non-invasive neuromodulation.
9. High-Resolution fMRI and Multivariate Representational Techniques
9.1 Ultra-High Field (7T) Imaging and Fine-Grained Topography
The evolution of magnetic resonance hardware from standard clinical 1.5T and 3T systems to ultra-high field 7 Tesla (7T) architectures fundamentally transformed the spatial scale of visual neuroscience. Ultra-high field imaging elevates the signal-to-noise ratio, allowing researchers to collect functional volumes with sub-millimeter isotropic voxel resolutions (e.g., $0.8 \text{ mm}^3$ or $0.65 \text{ mm}^3$). At this mesoscopic spatial scale, the internal micro-organization and laminar profiles of the human visual cortex can be interrogated directly in vivo.
Seven-Tesla fMRI investigations have provided unprecedented structural insight into the internal topography of the Face Fusiform Area. High-resolution mapping by researchers such as Federico De Martino, Kamil Ugurbil, and Kalanit Grill-Spector revealed that the FFA is not an undifferentiated mass of uniform tissue. Instead, sub-millimeter imaging definitively uncoupled the two distinct sub-regions, FFA-1 and FFA-2, showing that they sit along the fundus and lateral bank of the mid-fusiform sulcus, separated by an intervening strip of cortex with divergent connective and receptive field architectures.
Furthermore, population receptive field (pRF) mapping at 7T established that FFA-1 and FFA-2 display differing functional eccentricities. Receptive fields in FFA-1 are smaller and more retinotopically constrained to the central fovea, whereas FFA-2 features significantly larger, bilateral population receptive fields that integrate visual information across broader swathes of the visual field. Laminar fMRI has also begun to resolve depth-dependent BOLD responses across the cortical layers of the FFA, hinting at the segregation of feedforward sensory inputs in middle cortical layer IV from feedback predictive modulations concentrated within superficial and deep infragranular layers.
9.2 fMRI Adaptation (Repetition Suppression) Paradigms
Standard univariate fMRI is intrinsically constrained by the physical limits of the voxel: a typical $2 \text{ mm}^3$ imaging voxel contains millions of individual neurons, thousands of synapses, and complex local vascular networks. Consequently, if two distinct visual stimuli activate the same overall volume of tissue within the FFA with equal metabolic intensity, conventional subtraction analyses cannot determine whether those stimuli are being processed by the identical neural population or by intermingled, distinct sub-populations of cells.
To overcome this limitation, researchers adopted fMRI Adaptation (also known as Repetition Suppression), a paradigm pioneered in the visual domain by Kalanit Grill-Spector and Rafael Malach. fMRI adaptation exploits a universal neurophysiological property: when an identical visual stimulus or perceptual property is presented repeatedly, the specific sub-population of neurons tuned to that property exhibits a localized reduction in firing rate, which manifests as a sharp drop in the local BOLD signal. By systematically varying specific image dimensions across sequential presentations, investigators can probe the tuning properties of neural populations residing beneath the macroscopic voxel level.
When repetition suppression was applied to the FFA, the findings revealed the sophisticated computational nature of its representational code:
- Presenting the exact same face identity repeatedly produced massive BOLD adaptation in the FFA.
- If the face image was altered in low-level properties—such as changing its retinal size, shifting its visual position, or altering its illumination—the FFA continued to show profound adaptation, confirming that its neural populations encode abstract representations of identity that are invariant to simple physical changes.
- Conversely, when the identity was changed while holding lighting and expression constant, the adaptation was abolished, and the BOLD response rebounded completely.
- When faces were rotated in viewpoint (e.g., shifting from a front-facing angle to a 45-degree profile), anterior sub-patches of the FFA exhibited partial adaptation, demonstrating the emergence of view-tolerant neural tuning for individual humans.
9.3 Multivariate Pattern Analysis (MVPA) and Neural Decoding
In 2001, James Haxby and colleagues published an influential study in Science that presented a fundamental challenge to Kanwisher’s univariate modular framework. Haxby utilized Multivariate Pattern Analysis (MVPA) to show that even when the Face Fusiform Area was excluded from the analysis, the spatial pattern of voxel activations across the remainder of the ventral temporal cortex contained sufficient information to statistically decode whether a participant was viewing a face or an object. Furthermore, Haxby reported that the distributed pattern of activity within the FFA itself contained weak but decodable information about non-face categories (such as chairs or shoes). This sparked the “distributed vs. modular” representational debate, with Haxby arguing that category representations are intrinsically distributed over broad cortical swaths rather than locked inside modular boxes.
Kanwisher, working alongside researchers such as Hans Op de Beeck, Chris Baker, and Rebecca Saxe, responded by rigorously dissecting the conceptual difference between decodable information and computational utilization. While high-dimensional machine learning classifiers (such as Support Vector Machines) can exploit subtle, sub-threshold hemodynamic fluctuations across thousands of distributed voxels to achieve above-chance classification accuracy, this does not mean the brain’s downstream cognitive systems actually read out or rely upon that diffuse signal to guide perception. An information pattern can be a weak spatial correlation without performing a causal computational function.
Subsequent MVPA studies utilizing Representational Similarity Analysis (RSA), formalized by Nikolaus Kriegeskorte, succeeded in reconciling these paradigms. When high-resolution multivariate pattern classifiers were trained on the localized FFA, researchers discovered that its internal voxel patterns could reliably decode subtle individual face identities and discriminate between fine-grained variations in facial morphology. The consensus solidified: while distributed visual cortex exhibits faint residual cross-category information patterns, the FFA represents the specialized, high-fidelity computational core whose internal representational geometries are optimized for individual facial identification.
10. Developmental Trajectory and Plasticity of the FFA
10.1 Ontogeny of the FFA: Pediatric fMRI Studies
A central question in cognitive science is whether the functional modularity observed in the adult brain is present from birth, or whether it emerges through a protracted process of activity-dependent cortical development. To chart the ontogenetic trajectory of the Face Fusiform Area, pediatric fMRI investigations were initiated in the mid-2000s by researchers including Kalanit Grill-Spector, Golijeh Golarai, and K. Suzanne Scherf.
These pediatric studies, scanning children ranging from ages 4 to 12, as well as adolescents and adults, yielded striking insights into ventral stream development:
- In young children (ages 5 to 8), an independent functional localizer readily identifies face-selective voxels within the mid-fusiform gyrus, indicating that the basic topological coordinates of the FFA are established early in life.
- However, the total functional volume and spatial selectivity of the FFA undergo a prolonged maturation process. In young children, the FFA is significantly smaller in volume—frequently measuring only a fraction of its adult extent—and displays higher cross-responsiveness to non-face categories.
- Throughout late childhood and adolescence, the volume of the FFA expands substantially, carving out territory along the mid-fusiform sulcus until reaching adult dimensions.
This prolonged developmental expansion stands in stark contrast to the development of other ventral stream modules. For instance, the adjacent Parahippocampal Place Area (PPA) achieves adult-like volume and selectivity much earlier in childhood. Critically, this anatomical expansion of the FFA directly correlates with behavioral performance: as the volume and functional selectivity of the right FFA grow across childhood, children show parallel improvements in their behavioral memory for faces and their sensitivity to second-order relational spacing, demonstrating that the structural maturation of this cortical module underpins the acquisition of adult-level facial processing.
10.2 Innate Biases vs. Experiential Scaffolding
The protracted maturation of the FFA raises a profound question: does the initial specialization of the fusiform cortex depend entirely on visual experience, or does it originate from an innate, genetically specified neurobiological blueprint? Behavioral studies conducted with human neonates—most famously by Mark Johnson and John Morton in 1991—demonstrated that within minutes after birth, human infants display an innate orienting bias toward simple, high-contrast schematic configurations displaying a face-like geometry (three dark spots arranged in an inverted triangle representing eyes and a mouth; the “Conspec” mechanism).
This early innate bias is believed to be mediated by subcortical pathways encompassing the superior colliculus, the pulvinar nucleus of the thalamus, and the amygdala. This subcortical system automatically drives the infant’s foveal gaze toward conspecific faces in the immediate environment. By repeatedly foveating faces during early infancy, the developing brain projects an avalanche of high-resolution, centrally focused visual inputs directly into the retinotopically foveal-biased territory of the lateral fusiform gyrus. In this manner, an innate subcortical behavioral bias scaffolds the activity-dependent cortical specialization of the neocortex.
The definitive test of the necessity of early visual experience was achieved through non-human primate experiments conducted by Margaret Livingstone and Michael Arcaro in 2017. Macaque monkeys were reared from birth without ever seeing a human or monkey face (caregivers wore neutral, seamless face-obscuring masks). When these face-deprived macaques reached maturity and were scanned using fMRI, they completely failed to develop face patches in their inferotemporal cortex! Instead, their ventral streams developed specialized patches for other visual categories they had seen frequently, such as human hands. This demonstrated that while the cortical real estate of the lateral fusiform gyrus is predisposed to become the FFA due to its underlying connectivity and foveal retinotopic bias, the physical formation of the face module requires visual input during a sensitive critical period in early postnatal life.
10.3 Cross-Modal Plasticity and Tactile Face Recognition
Does the functional identity of the FFA depend exclusively on ocular visual signals, or does the region represent a more abstract, category-specific module capable of functioning across sensory modalities? This question led researchers to scan congenitally blind individuals—people who have never experienced optical visual inputs from birth—using high-resolution fMRI while they explored physical objects via the sense of touch.
Pioneering investigations by Pietro Pietrini, Marina Bedny, and Ella Striem-Amit examined whether tactile exploration of three-dimensional miniature faces, shoes, bottles, and geometric shapes could recruit the ventral occipitotemporal cortex in blind participants. The findings were remarkable: when congenitally blind adults haptically explored three-dimensional scale models of human faces, blood-oxygen-level-dependent activations were detected in the exact anatomical territory of the fusiform gyrus corresponding to the classical FFA. Tactile exploration of inanimate objects, in contrast, recruited lateral occipital regions corresponding to the lateral occipital tactile complex (LOtv).
These cross-modal plasticity discoveries carry profound epistemological implications for cognitive architecture. They suggest that the human brain does not simply organize its associative cortices on the basis of sensory input channels (vision, audition, touch). Instead, the cortex is organized according to abstract computational domains. The FFA may ultimately be an abstract “social entity identification module” whose native, default input in sighted individuals is optical vision, but whose underlying computational machinery is capable of repurposing its connectivity to process identity through tactile exploration when vision is absent.
11. Comparative Neuroscience: Homologous Networks in Non-Human Primates
11.1 Tsao and Freiwald’s Macaque Face-Patch System
The modern era of primate visual neuroscience witnessed an extraordinary empirical synthesis when Nancy Kanwisher’s human fMRI paradigms were adapted for non-human primates. In 2006, Doris Tsao, Winrich Freiwald, and Margaret Livingstone deployed functional neuroimaging in awake, fixating macaque monkeys (Macaca mulatta). By contrasting hemodynamic activations elicited by macaque and human faces against those elicited by common objects, toys, and abstract patterns, they discovered an interconnected network of discrete face-selective regions—the macaque “face-patch system”—distributed throughout the temporal lobe.
This macaque network comprises six distinct functional patches spanning the superior temporal sulcus and inferotemporal cortex:
- Posterior Lateral (PL): The most caudal patch along the inferior temporal cortex.
- Middle Lateral (ML) & Middle Fundus (MF): Centrally situated patches localized to the middle temporal gyrus and sulcal fundus.
- Anterior Lateral (AL) & Anterior Fundus (AF): Rostrally positioned processing hubs.
- Anterior Medial (AM): The most anterior patch, positioned near the temporal pole.
Comparative functional anatomy quickly established compelling homologies between the primate face patches and the human face network. The macaque middle patches (ML and MF) correspond functionally and topographically to the human Occipital Face Area (OFA) and the posterior FFA (FFA-1), whereas the anterior patches (AL and AM) represent the macaque evolutionary homolog of the anterior FFA (FFA-2) and the human anterior temporal face patch (ATFP).
11.2 Electrophysiological Tuning and Causal Inactivation Studies
The true power of the macaque face-patch discovery lay in its capacity to unite non-invasive neuroimaging with targeted, high-density single-unit electrophysiology. Utilizing high-resolution monkey fMRI as a stereotaxic surgical roadmap, Tsao and Freiwald lowered microelectrodes directly into the heart of the fMRI-defined middle face patches (ML/MF). The empirical results were staggering: across hundreds of recorded single units, an astonishing 90% to 97% of all individual neurons within the patch responded selectively to faces, exhibiting near-zero firing rates for any non-face object category.
This single-unit electrophysiology provided the ultimate biological validation of Kanwisher’s original modularity hypothesis. It proved that the elevated BOLD signals observed in the human FFA were not the averaged sum of intermingled, generic visual cells; rather, they reflected the collective metabolic expenditure of a dense, nearly pure population of face-tuned neurons packed into a dedicated functional patch. Furthermore, subsequent single-unit interrogations revealed the stepwise transformation of the neural code across the network:
- Neurons in patches ML/MF exhibited view-dependent tuning, firing vigorously to faces presented at specific angles.
- Neurons in anterior patch AL computed mirror-symmetric representations (firing identically to a face turned 45 degrees to the left or 45 degrees to the right).
- Neurons in the most anterior patch AM achieved full view-invariance, encoding the abstract individual identity of a face regardless of orientation, lighting, or expression.
Causal confirmation was achieved via targeted pharmacological inactivation experiments conducted by researchers such as Arash Afraz, James DiCarlo, and Winrich Freiwald. Micro-infusions of the GABA agonist muscimol directly into the macaque face patches produced immediate, reversible, and category-selective deficits in the animals’ ability to discriminate individual faces, while leaving object and color discrimination completely unperturbed. These causal interventions definitively linked single-neuron selectivity to conscious behavioral performance.
11.3 Evolutionary Origins of Specialized Neural Hardware for Social Cognition
The presence of an exquisitely conserved face-patch system across both Old World primates (macaques) and hominids (humans) indicates that this dedicated neuroanatomical hardware possesses deep evolutionary antiquity, tracing back to a common ancestor that lived at least 25 to 30 million years ago. What evolutionary pressures drove the biological emergence of such extreme cortical modularity?
The primary evolutionary driver was almost certainly the intricate demands of primate social cognition. Primates are among the most intensely social mammals on Earth, living within complex hierarchies, coalitions, and kinship structures. In such an environment, the fitness consequences of misidentifying a dominant conspecific, failing to recognize a kin relation, or misinterpreting a subtle facial threat or submission display are immediate and catastrophic. Rapid, automatic, and computationally efficient face recognition was a life-or-death imperative.
Building dedicated, modular cortical circuits allowed the primate brain to achieve three critical computational efficiencies:
- Wiring Minimization: Grouping millions of neurons with shared computational tuning into a compact, localized patch minimized axonal wire length, cutting down conduction latency and energetic metabolic costs.
- Dedicated Algorithmic Optimization: It permitted local microcircuits to implement non-linear, holistic geometric computations tailored specifically to the unique configuration of faces without interfering with the visual algorithms needed to identify tools, fruit, or predators.
- Rapid Processing Streams: It established direct, prioritized anatomical highways routing social identity data straight to the amygdala, anterior temporal semantic networks, and prefrontal decision nodes.
12. The Epistemological Legacy and Contemporary Impact of Kanwisher’s Work
12.1 Transformation of Neuroimaging Methodological Standards
The methodological legacy of Nancy Kanwisher’s 1997 experiments extends far beyond the study of face perception. The individual-subject functional Region of Interest (fROI) localizer paradigm that Kanwisher invented fundamentally transformed the analytical standards of modern cognitive neuroscience. Prior to her work, the field was dominated by group-averaging approaches that routinely obscured functional boundaries. Following the success of the FFA paradigm, the fROI method was rapidly adopted across every major domain of human cognitive neuroimaging.
Researchers worldwide deployed functional localizers to discover and map an entire constellation of dedicated functional components across the human neocortex:
- The Visual Word Form Area (VWFA) in the left fusiform gyrus (Laurent Cohen & Stanislas Dehaene).
- The Parahippocampal Place Area (PPA) for environmental scenes (Russell Epstein & Nancy Kanwisher).
- The Extrastriate Body Area (EBA) for human bodies (Paul Downing & Nancy Kanwisher).
- The Temporoparietal Junction (TPJ) module dedicated to Theory of Mind and mentalizing (Rebecca Saxe & Nancy Kanwisher).
- The frontotemporal language system defined via individual language localizers (Evelina Fedorenko & Nancy Kanwisher).
Furthermore, Kanwisher’s rigorous separation of localizer datasets from experimental evaluation datasets laid the conceptual groundwork for the modern replication crisis interventions in neuroimaging. Her framework established strict mathematical firewalls preventing circular statistical analyses (the “voodoo correlations” and double-dipping fallacies formalized by Nikolaus Kriegeskorte and Edward Vul in the late 2000s). By demanding that functional hypotheses be evaluated only on independent data partitions, Kanwisher instituted an empirical standard that made fMRI a far more robust, replicable, and respected scientific methodology.
12.2 Contributions to the Cognitive Architecture of the Human Mind
On a theoretical level, Kanwisher’s documentation of the FFA delivered a powerful empirical victory for modular theories of the human mind. The rise of connectionist artificial intelligence and neural network modeling in the late twentieth century had led many cognitive scientists to view the neocortex as an undifferentiated, general-purpose learning machine whose apparent functional specializations were merely arbitrary, transient ripples in a continuous, distributed weight matrix. Kanwisher’s empirical discoveries showed that this extreme domain-general view was fundamentally incomplete.
Her work established that the adult human brain is a structured organ composed of a dual architecture: it features an array of highly specialized, domain-specific computational processors (such as the FFA, PPA, and VWFA) dedicated to solving evolutionarily conserved or culturally intensive problems, embedded alongside powerful, flexible, domain-general processing networks (such as the frontoparietal “multiple-demand” system mapped by John Duncan) that handle novel, ad-hoc cognitive operations. This architectural dichotomy provided an empirical bridge reuniting Jerry Fodor’s philosophical modularity, Noam Chomsky’s universal grammar intuitions, and cognitive psychology’s information-processing frameworks with contemporary neurobiology.
Furthermore, the FFA has become the premier biological benchmark for modern computational neuroscience and artificial intelligence. In recent years, deep convolutional neural networks (CNNs) trained on real-world computer vision tasks (such as ImageNet) have been scrutinized by researchers like Dan Yamins and James DiCarlo. When deep networks are trained to optimize both general object recognition and face identification simultaneously, the artificial networks spontaneously bifurcate their internal layers, segregating their final representational layers into face-dedicated and object-dedicated branches that mimic the anatomical and functional separation between the human FFA and the lateral occipital complex. The FFA is no longer an anatomical mystery; it is understood as a mathematically optimal computational solution to the challenge of high-dimensional visual discrimination.
12.3 Open Frontiers and Lingering Questions in FFA Research
Despite more than a quarter-century of intensive empirical investigation, research into the Face Fusiform Area remains a dynamic and evolving frontier in cognitive neuroscience. As neuroimaging tools transition from macroscopic voxel averaging to microscopic, circuit-level interrogation, several fundamental questions remain unanswered:
First, visual neuroscientists are working to decipher the exact microcircuit-level computations that occur across the individual laminar layers of the FFA. With the deployment of ultra-high-field 7T and 9.4T scanners, researchers are now designing experiments to test how top-down predictive signals from the prefrontal cortex and amygdala integrate with bottom-up sensory evidence entering Layer IV of the fusiform cortex. These investigations aim to evaluate predictive coding models, testing whether the FFA primarily computes raw sensory representations or whether it dynamically processes “prediction errors”—the mathematical mismatch between what our social expectations anticipate and what our eyes physically observe.
Second, the scientific community is striving to understand the precise structural and functional connectome that routes information into and out of the FFA. Using advanced diffusion tractography, researchers are tracking the white matter highways—specifically the inferior longitudinal fasciculus (ILF) and the inferior fronto-occipital fasciculus (IFOF)—that link the FFA to anterior temporal memory networks, the social-emotional circuits of the amygdala, and the executive control nodes of the inferior frontal gyrus. Mapping these anatomical conduits is vital for understanding how the visual percept generated in the fusiform gyrus is converted into semantic identification, emotional recognition, and communicative action.
Finally, Nancy Kanwisher’s research enterprise continues to pursue a broader, breathtaking scientific question: what is the complete functional catalogue of the human brain? Through her ongoing public lectures, open-access datasets, and research initiatives at MIT, Kanwisher envisions a comprehensive functional atlas that maps every specialized computational component of the human mind. The Face Fusiform Area will forever stand as the first, foundational triumph in this scientific journey, proving that the deepest mysteries of human visual consciousness, social identity, and cognitive architecture can be unveiled through rigorous, empirical neuroimaging.
Conclusion
The discovery of the Face Fusiform Area by Nancy Kanwisher, Josh McDermott, and Marvin Chun in 1997 remains one of the definitive achievements in the history of cognitive neuroscience. By identifying a circumscribed patch of cortex within the human mid-fusiform gyrus that responds selectively to human faces, the landmark investigation resolved decades of clinical controversy surrounding acquired prosopagnosia and settled fierce philosophical disputes regarding the modular architecture of the mind. Kanwisher’s methodological insistence on the functional Region of Interest (fROI) localizer revolutionized neuroimaging analytical practices, establishing an empirical gold standard that rescued the discipline from the blurring artifacts of group averaging and circular statistical analyses.
Over the subsequent decades, an overwhelming convergence of empirical evidence—spanning invasive intracranial recordings, direct electrical brain stimulation in conscious patients, transcranial magnetic stimulation, pediatric ontogeny, cross-modal plasticity in the blind, non-human primate electrophysiology, and deep artificial neural networks—has corroborated Kanwisher’s domain-specific thesis. The FFA is not a generic engine for visual expertise, nor is it an epiphenomenal byproduct of distributed associative networks; it is a causally necessary, functionally specialized computational organ sculpted by natural selection and developmental experience to process the unique geometry of the human face.
Ultimately, the story of the Face Fusiform Area illuminates the foundational logic of the human brain. The mind is neither an undifferentiated, equipotential sponge nor an assembly of isolated, disconnected processors; it is a structured, elegant biological architecture that integrates specialized domain-specific modules with flexible, domain-general computational systems. By unmasking the neural locus where visual biology meets social recognition, Nancy Kanwisher’s pioneering experiments did not merely identify an anatomical structure—they provided a profound new window into the physical machinery that enables us to look into the eyes of another human being and recognize who they are.
References
- Allison, T., Puce, A., Spencer, D. D., & McCarthy, G. (1999). Electrophysiological studies of human face perception. I: Potentials generated in occipitotemporal cortex by faces and ball-bearing stimuli. Cerebral Cortex, 9(5), 415–430. https://doi.org/10.1093/cercor/9.5.415
- Arcaro, M. J., Schade, P. F., Vincent, J. L., Ponce, C. R., & Livingstone, M. S. (2017). Seeing faces is necessary for face-domain formation. Nature Neuroscience, 20(10), 1404–1412. https://doi.org/10.1038/nn.4635
- Bodamer, J. (1947). Die Prosop-Agnosie; die Agnosie des Physiognomieerkennens. Archiv für Psychiatrie und Nervenkrankheiten, 179(1–2), 6–53. https://doi.org/10.1007/BF00352849
- Damasio, A. R., Damasio, H., & Van Hoesen, G. W. (1982). Prosopagnosia: Anatomic basis and behavioral mechanisms. Neurology, 32(4), 331–341. https://doi.org/10.1212/wnl.32.4.331
- Downing, P. E., Jiang, Y., Shuman, M., & Kanwisher, N. (2001). A cortical area selective for visual processing of the human body. Science, 293(5539), 2470–2473. https://doi.org/10.1126/science.1063414
- Epstein, R., & Kanwisher, N. (1998). A cortical representation of the local visual environment. Nature, 392(6676), 598–601. https://doi.org/10.1038/33402
- Fodor, J. A. (1983). The Modularity of Mind: An Essay on Faculty Psychology. MIT Press. https://mitpress.mit.edu/9780262560252/the-modularity-of-mind/
- Gauthier, I., Tarr, M. J., Anderson, A. W., Skudlarski, P., & Gore, J. C. (1999). Activation of the middle fusiform ‘face area’ increases with expertise in recognizing novel objects. Nature Neuroscience, 2(6), 568–573. https://doi.org/10.1038/9224
- Golarai, G., Ghahremani, D. G., Whitfield-Gabrieli, S., Reiss, A., Eberhardt, J. L., Gabrieli, J. D. E., & Grill-Spector, K. (2007). Differential development of high-level visual cortex correlates with category-specific recognition memory. Nature Neuroscience, 10(4), 512–522. https://doi.org/10.1038/nn1865
- Grill-Spector, K., Knouf, N., & Kanwisher, N. (2004). The fusiform face area subserves face perception, not generic subordinate-level object categorization. Nature Neuroscience, 7(5), 555–562. https://doi.org/10.1038/nn1224
- Gross, C. G., Rocha-Miranda, C. E., & Bender, D. B. (1972). Visual properties of neurons in inferotemporal cortex of the macaque. Journal of Neurophysiology, 35(1), 96–111. https://doi.org/10.1152/jn.1972.35.1.96
- Haxby, J. V., Gobbini, M. I., Furey, M. L., Ishai, A., Schouten, J. L., & Pietrini, P. (2001). Distributed and overlapping representations of faces and objects in ventral temporal cortex. Science, 293(5539), 2425–2430. https://doi.org/10.1126/science.1063736
- Haxby, J. V., Hoffman, E. A., & Gobbini, M. I. (2000). The distributed human neural system for face perception. Trends in Cognitive Sciences, 4(6), 223–233. https://doi.org/10.1016/S1364-6613(00)01482-0
- Johnson, M. H., Dziurawiec, S., Ellis, H., & Morton, J. (1991). Newborns’ preferential tracking of face-like stimuli and its subsequent decline. Cognition, 40(1–2), 1–19. https://doi.org/10.1016/0010-0277(91)90045-6
- Kanwisher, N. (2000). Domain specificity in face perception. Nature Neuroscience, 3(8), 759–763. https://doi.org/10.1038/77664
- Kanwisher, N. (2010). Functional specificity in the human brain: A window into the functional architecture of the mind. Proceedings of the National Academy of Sciences, 107(25), 11163–11170. https://doi.org/10.1073/pnas.1005062107
- Kanwisher, N., McDermott, J., & Chun, M. M. (1997). The fusiform face area: A module in human extrastriate cortex specialized for face perception. The Journal of Neuroscience, 17(11), 4302–4311. https://doi.org/10.1523/JNEUROSCI.17-11-04302.1997
- Kriegeskorte, N., Mur, M., & Bandettini, P. A. (2008). Representational similarity analysis – connecting the branches of systems biology. Frontiers in Systems Neuroscience, 2, 4. https://doi.org/10.3389/neuro.06.004.2008
- Kwong, K. K., Belliveau, J. W., Chesler, D. A., Goldberg, I. E., Weisskoff, R. M., Poncelet, B. P., Kennedy, D. N., Hoppel, B. E., Cohen, M. S., & Turner, R. (1992). Dynamic magnetic resonance imaging of human brain activity during primary sensory stimulation. Proceedings of the National Academy of Sciences, 89(12), 5675–5679. https://doi.org/10.1073/pnas.89.12.5675
- Malach, R., Reppas, J. B., Benson, R. R., Kwong, K. K., Jiang, H., Kennedy, W. A., Ledden, P. J., Brady, T. J., Rosen, B. R., & Tootell, R. B. (1995). Object-related activity revealed by functional magnetic resonance imaging in human occipital cortex. Proceedings of the National Academy of Sciences, 92(18), 8135–8139. https://doi.org/10.1073/pnas.92.18.8135
- Ogawa, S., Lee, T. M., Kay, A. R., & Tank, D. W. (1990). Brain magnetic resonance imaging with contrast dependent on blood oxygenation. Proceedings of the National Academy of Sciences, 87(24), 9868–9872. https://doi.org/10.1073/pnas.87.24.9868
- Parvizi, J., Jacques, C., Foster, B. L., Withoft, N., Rangarajan, V., Weiner, K. S., & Grill-Spector, K. (2012). Electrical stimulation of human fusiform face-selective regions distorts face perception. The Journal of Neuroscience, 32(43), 14915–14920. https://doi.org/10.1523/JNEUROSCI.2609-12.2012
- Perrett, D. I., Rolls, E. T., & Caan, W. (1982). Visual neurones responsive to faces in the monkey temporal cortex. Experimental Brain Research, 47(3), 329–342. https://doi.org/10.1007/BF00239352
- Pietrini, P., Furey, M. L., Ricciardi, E., Gobbini, M. I., Wu, W. H., Cohen, L., Le Bihan, D., & Haxby, J. V. (2004). Beyond sensory images: Object-based representation in the human ventral pathway. Proceedings of the National Academy of Sciences, 101(15), 5658–5663. https://doi.org/10.1073/pnas.0400858101
- Pitcher, D., Walsh, V., Yovel, G., & Duchaine, B. (2007). TMS evidence for the involvement of the right occipital face area in early face processing. Current Biology, 17(18), 1568–1573. https://doi.org/10.1016/j.cub.2007.07.063
- Saxe, R., & Kanwisher, N. (2003). People thinking about thinking people: The role of the temporo-parietal junction in “theory of mind”. NeuroImage, 19(4), 1835–1842. https://doi.org/10.1016/S1053-8119(03)00230-1
- Scherf, K. S., Behrmann, M., Humphreys, K., & Luna, B. (2007). Visual category-selectivity for faces, places and objects emerges along different developmental trajectories. Developmental Science, 10(4), F15–F30. https://doi.org/10.1111/j.1467-7687.2007.00595.x
- Sergent, J., Ohta, S., & MacDonald, B. (1992). Functional neuroanatomy of face and object processing: A positron emission tomography study. Brain, 115(1), 15–36. https://doi.org/10.1093/brain/115.1.15
- Tsao, D. Y., Freiwald, W. A., Tootell, R. B., & Livingstone, M. S. (2006). A cortical region consisting entirely of face-selective cells. Science, 311(5761), 670–674. https://doi.org/10.1126/science.1119983
- Weiner, K. S., Golarai, G., Caspers, J., Chu, D., Fink, C., Malach, R., & Grill-Spector, K. (2014). The mid-fusiform sulcus: An anatomical landmark for visual category-selective areas in human high-level visual cortex. NeuroImage, 84, 453–469. https://doi.org/10.1016/j.neuroimage.2013.08.068
- Yin, R. K. (1969). Looking at upside-down faces. Journal of Experimental Psychology, 81(1), 141–145. https://doi.org/10.1037/h0027474