Abstract
The Need for Evaluation (NFE) scale is a widely utilized psychometric instrument developed by Richard E. Petty and W. Blair G. Jarvis (1996) to operationalize individual differences in the chronic propensity to engage in evaluative responding. While classical social psychological models traditionally assumed that evaluation is a universal and spontaneous reaction to encountering any stimulus object, modern dispositional research demonstrates substantial inter-individual variability in the habitual tendency to judge entities as good or bad, favorable or unfavorable, and desirable or undesirable. The standardized instrument comprises 16 self-report items evaluated on a 5-point Likert response format ranging from 1 (“extremely uncharacteristic of you”) to 5 (“extremely characteristic of you”), with alternative administration formats utilizing strongly disagree to strongly agree anchors. Extensive psychometric testing confirms a predominant single-factor structure accounting for broad evaluative drive, supported by high internal consistency reliability (typically ranging from Cronbach’s alpha = .82 to .88) and high temporal stability across test-retest intervals ($r = .83$ over 10 weeks). Construct validation studies demonstrate that high-NFE individuals naturally hold more attitudes across diverse socio-political, consumer, and everyday topics, access their stored evaluations more rapidly in cognitive reaction-time paradigms, experience lower rates of “no opinion” or neutral responses on survey measures, and exhibit heightened susceptibility to framing effects that emphasize evaluative dimensions. The scale has demonstrated strong convergent validity with related cognitive constructs, such as the Need for Cognition, while maintaining clear discriminant validity from general cognitive ability, neuroticism, extraversion, and social desirability response biases. This comprehensive psychometric review explores the theoretical foundations, structural factor characteristics, cross-cultural validity, and methodological implementation parameters of the NFE scale across social, political, and consumer psychology research contexts.
Keywords
Need for Evaluation, NFE, Attitude Formation, Evaluative Responding, Social Cognition, Psychometrics, Scale Validation, Individual Differences, Richard Petty, Attitude Accessibility
Authors
The Need for Evaluation (NFE) construct and its primary measurement tool were conceptualized and psychometrically validated by W. Blair G. Jarvis and Richard E. Petty at the Department of Psychology, The Ohio State University (Columbus, Ohio, United States). Dr. Richard E. Petty is a Distinguished University Professor Emeritus of Psychology, internationally recognized for foundational contributions to attitudes and social cognition research, including the co-development of the Elaboration Likelihood Model (ELM). Dr. W. Blair G. Jarvis contributed extensively to the empirical differentiation of evaluative motivation from general cognitive motivation during his doctoral and postdoctoral tenure at The Ohio State University. In psychometric and survey methodology literature, historical and comparative validation studies also assess the scale against classic dispositional metrics, such as the Marlowe-Crowne Social Desirability Scale by Douglas P. Crowne and David Marlowe (1960), ensuring that evaluative responsiveness is not an artifact of impression management or psychopathological response tendencies.
Purpose
The primary purpose of the Need for Evaluation (NFE) scale is to measure an individual’s chronic tendency to form, hold, and express evaluative judgments across diverse life domains, environmental stimuli, and conceptual issues. Evaluative responding—categorizing objects, people, behaviors, and ideas along a positive-to-negative valence continuum—is one of the most fundamental cognitive operations performed by the human mind. However, empirical work in social cognition revealed that human beings do not engage in evaluation with uniform frequency or intensity. While some individuals spontaneously and systematically assess almost every environmental stimulus they encounter as either good or bad, others remain relatively neutral, forming evaluations only when situational demands explicitly require them to do so.
In academic research, the NFE scale serves to isolate this intrinsic evaluative motive from related but distinct constructs, such as the volume of intellectual effort an individual enjoys (Need for Cognition) or the discomfort experienced when facing ambiguity or incomplete information (Need for Closure). By administering the scale, behavioral scientists can test hypotheses regarding why certain individuals hold pre-formed, highly accessible attitudes on obscure political issues, cultural phenomena, and consumer products, whereas others consistently select “don’t know” or midpoint options in survey research.
In applied contexts, such as political polling, market research, and communication science, the NFE scale serves several vital diagnostic and predictive functions:
- Predicting Opinion Formation and Polling Behavior: High-NFE citizens naturally crystallize political viewpoints, exhibit lower non-response rates on public policy surveys, and show marked stability in their voting choices over campaign cycles. Conversely, low-NFE voters are more likely to make late-breaking decisions and remain uncommitted until proximal external cues force a selection.
- Consumer Decision-Making Analysis: In consumer behavior studies, high-NFE individuals spontaneously form brand preferences, generate faster product appraisals, and maintain definitive hierarchies of likes and dislikes. Low-NFE consumers frequently evaluate products strictly according to situational task parameters rather than enduring evaluative dispositions.
- Clinical and Counseling Psychology Applications: While not a clinical diagnostic instrument for psychopathology, the NFE scale aids clinicians in understanding cognitive styles characterized by hyper-judgmentalism, rigid dichotomous thinking, or, at the other extreme, excessive ambivalence and decision-paralysis. Individuals who score extremely high on NFE may exhibit a tendency toward black-and-white thinking that exacerbates interpersonal conflict, perfectionism, or evaluative anxiety.
Psychological Construct
The psychological construct captured by the Need for Evaluation scale represents an enduring, trait-like individual difference variable situated within social cognitive personality theory. It reflects the chronic intrinsic motivation to evaluate stimuli—determining whether objects, events, concepts, or persons are good versus bad, desirable versus undesirable, or pleasant versus unpleasant. The construct is conceptualized primarily as a continuous, unipolar dimension extending from low to high chronic evaluative drive.
Core Features of High Need for Evaluation
Individuals characterized by a high need for evaluation exhibit distinct cognitive, affective, and behavioral tendencies:
- Spontaneous Valence Assignment: High-NFE individuals habitually attach positive or negative labels to newly encountered stimuli without explicit external prompts or practical utility. For instance, when walking into an unfamiliar room or hearing an unfamiliar song, high-NFE individuals immediately determine whether they like or dislike the environment or music.
- Attitude Accessibility: Because evaluations are generated continuously, high-NFE individuals possess a rich reservoir of chronic, accessible attitudes stored in long-term memory. In cognitive latency paradigms, these individuals respond significantly faster to attitude queries because their evaluations are pre-computed rather than constructed on the spot.
- Discomfort with Evaluative Neutrality: Individuals high in NFE experience a motivational preference for clear, definite views and find fence-sitting or ambivalence uncomfortable. When presented with ambiguous social dilemmas, they actively strive to resolve evaluative uncertainty into definitive stances.
- Judgmental Breadth: High-NFE individuals form opinions not only about personally relevant matters but also about issues, objects, and abstract domains that do not directly affect their personal welfare (e.g., modern art genres, distant political conflicts, architectural styles).
Core Features of Low Need for Evaluation
Conversely, individuals with a low need for evaluation demonstrate a cognitive style that prioritizes descriptive or phenomenological processing over immediate valence assignment:
- Observational and Descriptive Focus: Low-NFE individuals are comfortable observing, categorizing, and describing stimuli without assigning evaluative tags. When walking through an art exhibit, they may observe the brushwork, color palette, and geometric composition without feeling a compulsion to decide whether the artwork is inherently “good” or “bad.”
- Attitude Construction on Demand: Rather than retrieving pre-existing evaluations, low-NFE individuals construct attitudes online only when situational contingencies demand an explicit choice or rating. Consequently, their reported attitudes exhibit greater contextual variability and are more heavily influenced by immediate context effects, question framing, and mood states.
- Comfort with Neutrality: These individuals readily accept neutrality, ambivalence, and non-commitment as valid and comfortable epistemic states. They report little distress when having “no opinion” on complex political, social, or aesthetic debates.
Theoretical Framework
The conceptual foundation of the Need for Evaluation scale emerges from functional attitude theories and contemporary models of social information processing, primarily spearheaded by functional theories of attitudes (e.g., Katz, 1960; Smith, Bruner, & White, 1956) and cognitive motivation paradigms (Petty & Cacioppo, 1986). Historically, social psychologists treated evaluation as an involuntary, inevitable cognitive reflex. Pioneering models of automatic attitude activation (e.g., Fazio, Sanbonmatsu, Powell, & Kardes, 1986) suggested that encountering any attitude object automatically triggers associated evaluations from memory.
However, Jarvis and Petty (1996) proposed an individual difference framework asserting that while the *capacity* for automatic evaluation may be universal, the *motivation* to continuously form and store valence tags varies substantially across the population. This perspective integrated three primary theoretical mechanisms:
1. The Object-Appraisal Function of Attitudes
Functional attitude theory posits that attitudes serve critical psychological functions, the most prominent being the “object-appraisal” or knowledge function. Stored evaluations act as cognitive shortcuts (heuristics) that allow organisms to rapidly categorize objects in the environment as threats to be avoided or opportunities to be approached. Jarvis and Petty argued that individuals vary in the value they place on this object-appraisal mechanism. For high-NFE individuals, maintaining a comprehensively evaluated cognitive map of the world is perceived as highly functional and intrinsically rewarding, enabling decisive navigation through complex environments.
2. Cognitive Miser vs. Epistemic Drive Models
Classic cognitive psychology frequently characterizes humans as “cognitive misers” who conserve mental energy by avoiding unnecessary computational effort unless motivated by external rewards or survival needs. The NFE model establishes that evaluation represents a distinct form of mental expenditure. While some cognitive misers minimize evaluative processing (low NFE), other individuals possess an epistemic disposition that finds evaluative categorization inherently satisfying and effortless. Crucially, Jarvis and Petty demonstrated that NFE is conceptually distinct from the Need for Cognition (NFC). Whereas NFC reflects the intrinsic desire to engage in effortful, analytic information processing and problem-solving (enjoying complex reasoning, puzzles, and intellectual depth), NFE reflects the specific motivation to categorize objects along a valence continuum (good vs. bad). An individual can be high in NFC and low in NFE (e.g., a scientist who loves dissecting the mechanical complexities of a system without caring to label it “good” or “bad”), or high in NFE and low in NFC (e.g., an individual who rapidly forms strong opinions on every topic based on gut feelings without examining complex underlying data).
3. The Dual-Process Perspective
Within dual-process architectures of persuasion and judgment, such as the Elaboration Likelihood Model, the Need for Evaluation functions as an individual difference variable that influences how arguments and messages are processed. High-NFE individuals are motivated to extract evaluative conclusions from persuasive communications regardless of situational cues, whereas low-NFE individuals require high personal relevance or explicit task prompts before elaborating on the evaluative implications of a message.
Validity
The psychometric validity of the Need for Evaluation scale has been confirmed across diverse laboratory experiments, field studies, and representative survey investigations.
Construct and Known-Groups Validity
Construct validity was established in the initial validation studies by Jarvis and Petty (1996) across multiple experimental designs. In one foundational paradigm, participants were exposed to a series of novel paintings and abstract stimuli. High-NFE participants spontaneously produced significantly more evaluative thoughts (e.g., “This painting is hideous,” “I love the balance of colors”) in open-ended thought-listing protocols than low-NFE participants, who predominantly listed descriptive thoughts (e.g., “The image contains blue squares”). Furthermore, when completing comprehensive lifestyle and public policy surveys, high-NFE individuals consistently selected the “no opinion” or neutral response alternatives at a significantly lower frequency ($p < .001$) compared to their low-NFE counterparts.
Convergent Validity
Convergent validity is supported by meaningful correlations with theoretically aligned constructs:
- Attitude Accessibility: Studies measuring response latencies reveal that high-NFE individuals exhibit faster reaction times when reporting their evaluations of social, political, and consumer objects, confirming that their attitudes are chronic and readily accessible in memory ($r = -.34$ to $-.42$ with response latencies).
- Need for Cognition: Meta-analyses and initial studies report a moderate positive correlation between NFE and the Need for Cognition (typically $r = .25$ to $.35$), reflecting an overlap in cognitive engagement while confirming that they measure distinct psychological phenomena.
- Personal Need for Structure and Need for Closure: NFE demonstrates weak-to-moderate positive correlations with the Personal Need for Structure ($r \approx .15$ to $.22$) and the Need for Cognitive Closure ($r \approx .18$), as evaluative closure helps organize one’s perceptual environment.
Discriminant Validity
Discriminant validity has been rigorously demonstrated against several major personality dimensions and response biases:
- Social Desirability: Correlations with the Marlowe-Crowne Social Desirability Scale (Crowne & Marlowe, 1960) and the Balanced Inventory of Desirable Responding (BIDR) are consistently negligible (ranging from $r = -.08$ to $r = .05$, non-significant), demonstrating that the scale does not tap into the desire to appear socially agreeable or conventional.
- Five-Factor Model of Personality: NFE exhibits low correlations with Extraversion ($r \approx .12$), Neuroticism ($r \approx -.05$), Agreeableness ($r \approx -.10$), and Conscientiousness ($r \approx .08$). A modest positive correlation is sometimes observed with Openness to Experience ($r \approx .15$ to $.20$).
- Cognitive Ability: Jarvis and Petty (1996) verified that NFE is unrelated to measures of verbal intelligence, American College Test (ACT) scores, and Grade Point Average (GPA), proving that the scale captures an evaluative motivation rather than intellectual aptitude.
Predictive and Ecological Validity
In political psychology, studies using American National Election Studies (ANES) data (e.g., Bizer et al., 2004) showed that voters high in NFE held more candidate attitudes, evaluated political figures more decisively, reported higher political participation, and were less susceptible to survey-wording artifacts and status-quo framing biases.
Reliability
The Need for Evaluation scale demonstrates robust internal consistency and temporal reliability across collegiate, community, and nationally representative adult samples.
Internal Consistency
Across the development samples reported by Jarvis and Petty (1996), the 16-item scale exhibited high internal consistency reliability:
- Initial Undergraduate Development Sample ($N = 464$): Cronbach’s alpha was computed at $\alpha = .87$.
- Cross-Validation Sample ($N = 398$): Cronbach’s alpha remained robust at $\alpha = .88$.
- Representative General Population Samples: Subsequent administrations in national surveys (such as the ANES pilot testing modules) reported Cronbach’s alpha coefficients ranging between $\alpha = .80$ and $\alpha = .85$.
- Item-Total Correlations: Corrected item-total correlations across the 16 items typically range from $.35$ to $.65$, with no single item’s deletion yielding an increase in the composite scale alpha.
Test-Retest Stability
To verify the temporal stability of the scale as a chronic dispositional trait, Jarvis and Petty (1996) administered the instrument across a 10-week test-retest interval to an undergraduate cohort ($N = 104$). The test-retest reliability coefficient was $r = .83$ ($p < .001$), confirming substantial temporal stability over extended periods and demonstrating that the NFE instrument measures an enduring personality orientation rather than a fluctuating affective state.
Factor Analysis
The latent structure of the Need for Evaluation scale has been evaluated using both Exploratory Factor Analysis (EFA) and Confirmatory Factor Analysis (CFA).
Exploratory Factor Analysis (EFA)
In the original psychometric construction of the scale, an initial pool of 46 items reflecting evaluative motivation was administered to participants. Principal Axis Factoring and Principal Components Analysis with both orthogonal (Varimax) and oblique (Promax) rotations were conducted. Scree plot analyses and eigenvalue evaluations revealed a dominant primary factor that accounted for the vast majority of common variance (eigenvalue > 5.0, explaining approximately 30% of total variance before item reduction). Items exhibiting factor loadings below .35, severe cross-loadings on secondary minor factors, or weak item-total correlations were iteratively eliminated, resulting in the final 16-item unidimensional instrument.
Confirmatory Factor Analysis (CFA) and Model Fit
Subsequent structural investigations have evaluated both a strict unidimensional model and a bifactor/two-factor model accounting for item valence (method effects associated with reverse-keyed items). Findings from CFA studies in literature (e.g., Jarvis & Petty, 1996; Bizer et al., 2004; Tormala & Petty, 2001) demonstrate the following structural parameters:
- Unidimensional Model Fit: When specifying a single latent Need for Evaluation factor, model fit indices generally meet acceptable psychometric thresholds: Comparative Fit Index (CFI) $= .90$ to $.94$, Tucker-Lewis Index (TLI) $= .89$ to $.93$, Root Mean Square Error of Approximation (RMSEA) $= .048$ to $.062$ ($90%\text{ CI } [.041, .068]$), and Standardized Root Mean Square Residual (SRMR) $= .042$ to $.055$.
- Method Effects for Negatively Keyed Items: In certain large-scale survey environments, a two-factor specification—separating directly keyed items (evaluative engagement) from reverse-keyed items (preference for neutrality/non-evaluation)—yields slightly elevated global fit indices (CFI > .95, RMSEA < .045). However, methodological analyses confirm that this two-factor solution is a statistical artifact of item phrasing directionality rather than substantively meaningful sub-traits. Consequently, psychometricians recommend treating the 16 items as a singular, unified construct.
- Factor Loadings: Standardized factor loadings ($lambda$) across the 16 items generally range between $.38$ and $.72$, with key anchoring items such as “I form opinions about everything” and “It is very important to me to hold strong opinions” exhibiting the strongest loadings on the general latent evaluative drive factor.
Instrument / Measurement Tool
The standardized instrument details and scoring protocols for the Need for Evaluation scale are structured as follows:
- Test Type: Self-report psychological personality assessment scale.
- Construct Assessed: Individual differences in chronic evaluative responding and attitude formation motivation.
- Item Count: 16 items.
- Administration Format: Paper-and-pencil or computerized self-administered questionnaire.
- Typical Completion Time: 3 to 5 minutes.
- Standard Response Scale: 5-point Likert scale (typically: 1 = extremely uncharacteristic of you, 2 = somewhat uncharacteristic of you, 3 = uncertain, 4 = somewhat characteristic of you, 5 = extremely characteristic of you; or 1 = strongly disagree to 5 = strongly agree).
- Scoring and Transformation Rules:
- The instrument consists of both directly scored and reverse-scored items.
- Reverse-Scored Items: Items 1, 3, 6, 8, 9, 11, and 12 must be reverse-coded prior to computing the aggregate score (i.e., $1 \rightarrow 5$, $2 \rightarrow 4$, $3 \rightarrow 3$, $4 \rightarrow 2$, $5 \rightarrow 1$) according to the specific reverse-scoring rules defined for this standardized protocol.
- Total Score Calculation: Calculate the sum or the arithmetic mean across all 16 items after completing necessary reverse coding.
- Score Range: Total summed scores range from 16 to 80 (or mean scores from 1.0 to 5.0). Higher composite scores reflect a stronger chronic dispositional need for evaluation, whereas lower scores reflect a relative disinclination to evaluate stimuli and comfort with evaluative neutrality.
Permissions & Fee and Test Year
The Need for Evaluation scale was originally published in 1996 by W. Blair G. Jarvis and Richard E. Petty in the Journal of Personality and Social Psychology, a flagship journal of the American Psychological Association (APA). As is standard for academic psychometric instruments published in scholarly journals, the scale is generally accessible without royalty fees for non-commercial scientific research, academic dissertations, and educational training, provided that appropriate formal bibliographic citation is given to the original authors and the American Psychological Association. For commercial testing, clinical diagnostics, or inclusion within proprietary corporate assessment platforms, permissions must be formally requested from the copyright holder (the American Psychological Association). No standard per-administration test fee is charged for independent academic research applications.
References
- Bizer, G. Y., Krosnick, J. A., Holbrook, A. L., Christian, S., Wheeler, S. C., & Petty, R. E. (2004). The need to evaluate: Differences in the formation and expression of political opinions. Journal of Personality, 72(5), 995–1024. https://doi.org/10.1111/j.1467-6494.2004.00288.x
- Crowne, D. P., & Marlowe, D. (1960). A new scale of social desirability independent of psychopathology. Journal of Consulting Psychology, 24(4), 349–354. https://doi.org/10.1037/h0047358
- Fazio, R. H., Sanbonmatsu, D. M., Powell, M. C., & Kardes, F. R. (1986). On the automatic activation of attitudes. Journal of Personality and Social Psychology, 50(2), 229–238. https://doi.org/10.1037/0022-3514.50.2.229
- Hermans, D., De Houwer, J., & Eelen, P. (2001). A time course analysis of the affective priming effect. Cognition & Emotion, 15(2), 143–165. https://doi.org/10.1080/02699930125768
- Jarvis, W. B. G., & Petty, R. E. (1996). The need to evaluate. Journal of Personality and Social Psychology, 70(1), 172–194. https://doi.org/10.1037/0022-3514.70.1.172
- Katz, D. (1960). The functional approach to the study of attitudes. Public Opinion Quarterly, 24(2), 163–204. https://doi.org/10.1086/266945
- Petty, R. E., & Cacioppo, J. T. (1986). The elaboration likelihood model of persuasion. Advances in Experimental Social Psychology, 19, 123–205. https://doi.org/10.1016/S0065-2601(08)60214-2
- Smith, M. B., Bruner, J. S., & White, R. W. (1956). Opinions and personality. John Wiley & Sons.
- Tormala, F. L., & Petty, R. E. (2001). On-line versus memory-based processing: The role of need to evaluate in person perception. Personality and Social Psychology Bulletin, 27(12), 1599–1612. https://doi.org/10.1177/01461672012712004
Items of the Scale
Instructions: For each of the following statements, please indicate how characteristic it is of you by using the response scale provided below.
Response Scale: 5-point Likert scale (typically: 1 = extremely uncharacteristic of you, 2 = somewhat uncharacteristic of you, 3 = uncertain, 4 = somewhat characteristic of you, 5 = extremely characteristic of you; or 1 = strongly disagree to 5 = strongly agree)
Reverse Scoring: Items 1, 3, 6, 8, 9, 11, and 12 are reverse-scored.
- I form opinions about everything.
- I prefer to avoid taking extreme positions.
- It is very important to me to hold strong opinions.
- I want to know exactly what is going on around me.
- I don’t like to have to make a lot of choices everyday.
- I often prefer to remain neutral on complex issues.
- If something does not affect me, I do not usually expend much effort thinking about it.
- I enjoy having clear and definite views on almost everything.
- I would rather decide what is right or wrong for myself than let other people decide for me.
- I like to have strong opinions even when I am not personally involved.
- I have many more opinions than the average person.
- I would rather not know the opinions of other people.
- I pay a lot of attention to whether I like or dislike things.
- I only form opinions about issues when I really have to.
- I am often hesitant to express my opinions.
- It bothers me when other people do not have opinions.