Abstract
The Evaluation Similarity (ES) scale is a psychometric instrument developed by consumer behavior and marketing researchers Andrew D. Gershoff, Ashesh Mukherjee, and Anirban Mukhopadhyay (2007) to measure the degree to which an individual perceives another entity—such as a peer, professional critic, human agent, or automated recommender system—as possessing similar evaluative criteria, tastes, and judgment standards. Initially introduced in the context of agent evaluation and social judgment research within the Journal of Consumer Research, the instrument was formulated to capture perceived alignment in aesthetic, functional, and subjective evaluation processes. The scale comprises three items administered via a multi-point Likert-type or semantic differential response format, assessing shared taste, convergent judgment strategies, and concordance in product appraisal. Across empirical investigations, the scale has exhibited robust psychometric properties, consistently demonstrating high internal consistency reliability (with Cronbach's alpha coefficients typically exceeding .85 and often surpassing .90) and solid construct validity. Confirmatory factor analyses across experimental conditions support a unidimensional structure that is invariant across varied product categories, including hedonic goods (e.g., films, music, experiential dining) and utilitarian offerings. By quantifying perceived judgment alignment, the Evaluation Similarity scale serves as an indispensable tool for marketing scientists, social psychologists, human-computer interaction (HCI) researchers, and behavioral decision theorists exploring word-of-mouth (WOM) transmission, algorithmic trust, advice utilization, and the interpersonal dynamics of preference attribution.
Keywords
Evaluation Similarity, Perceived Taste Similarity, Agent Evaluation, Consumer Judgment, Social Cognition, Recommendation Systems, Interpersonal Congruence, Psychometrics, Decision Making, Construct Validity
Authors
The Evaluation Similarity scale was conceptualized, operationalized, and validated by a team of prominent consumer psychologists and marketing scholars:
- Andrew D. Gershoff, Ph.D. — Professor of Marketing and Foley’s Department Chair in Retailing at the McCombs School of Business, The University of Texas at Austin. Dr. Gershoff’s research focuses on consumer decision-making, interpersonal evaluation, word-of-mouth, trust, and how individuals predict others’ tastes and infer others’ motives.
- Ashesh Mukherjee, Ph.D. — Associate Professor of Marketing at the Desautels Faculty of Management, McGill University. Dr. Mukherjee specializes in consumer information processing, digital marketing, the psychology of pricing, and interpersonal advice acceptance.
- Anirban Mukhopadhyay, Ph.D. — Lifestyle International Professor of Business and Chair Professor of Marketing at the Hong Kong University of Science and Technology (HKUST). Dr. Mukhopadhyay investigates self-regulation, consumer lay theories, interpersonal influence, and social judgment processes.
Purpose
The primary purpose of the Evaluation Similarity (ES) scale is to measure an individual’s subjective assessment of the degree to which another evaluator—termed an “agent”—shares their foundational standards of taste, critical values, and qualitative judgment. In social exchange and consumer contexts, decision-makers frequently face incomplete information when choosing among alternatives, leading them to rely on surrogates, critics, algorithmic collaborative filters, or personal acquaintances. However, advice acceptance and recommendation adoption are not merely functions of whether an agent expresses a positive or negative valence; they depend critically on whether the decision-maker perceives that agent as evaluating the consumption domain through a congruent lens.
From a theoretical standpoint, Gershoff, Mukherjee, and Mukhopadhyay (2007) constructed this instrument to isolate the specific mechanism through which agreement or disagreement on individual attributes influences global perceptions of an agent’s evaluative competence and taste alignment. In experimental consumer psychology, researchers frequently manipulate prior agreement (e.g., discovering that an agent loves or hates the same film as the participant) to observe downstream effects on advice discounting, future choice prediction, and trust. The ES scale provides an operational tool to confirm whether experimental manipulations successfully altered the participant’s mental representation of the agent’s taste profile, and whether this perceived similarity mediates ultimate choices.
Beyond academic laboratory experiments, the scale addresses critical challenges in applied settings, including:
- Algorithmic Recommendation and Human-AI Interaction: Assessing user trust in personalized recommender engines (e.g., Netflix, Spotify, Amazon). The scale evaluates whether users perceive algorithmic recommendation outputs as mirroring their personal taste profile or relying on discordant evaluative heuristics.
- Influencer Marketing and Word-of-Mouth (WOM): Determining the efficacy of opinion leaders. When followers perceive an influencer to possess high evaluation similarity, positive and negative endorsements carry greater diagnostic value and exert stronger persuasive impact.
- Interpersonal Decision Consultations: Investigating joint decision-making in couples, familial units, and organizational purchasing committees, where perceived evaluation similarity moderates consensus building, negotiation strategies, and delegation of decision authority.
Psychological Construct
The Evaluation Similarity scale operationalizes a focal facet of social perception and interpersonal congruence: the psychological attribution of shared evaluative schemas. Within psychometrics and behavioral decision research, perceived similarity is often decomposed into demographic similarity, value similarity, and taste/evaluative similarity. Evaluation similarity specifically concerns the cognitive and affective appraisal processes an entity employs when judging a target stimulus.
The construct encompasses three tightly interrelated conceptual facets:
- Shared Taste: The experiential alignment of subjective preferences. In hedonic domains such as art, entertainment, sensory foods, and aesthetics, taste represents an internal hedonic response. Attributing shared taste to an agent means assuming that identical sensory or narrative stimuli will elicit comparable hedonic reactions in both the self and the agent.
- Congruent Evaluative Standards: The criteria, weights, and metrics applied during judgment. Two individuals might both appreciate a product, but one may prioritize durability and functional efficiency while the other values aesthetic elegance and novelty. Evaluation similarity reflects the belief that the agent uses the same underlying attribute weighting schemes and threshold standards as the focal judge.
- Predictive Judgment Equivalence: The functional outcome of the evaluative process. This dimension captures the expectation that when presented with an unencountered object or choice set within the specified domain, the agent will render verdicts, ratings, or selections highly congruent with those the respondent would reach independently.
In the seminal framework of Gershoff et al. (2007), evaluation similarity is distinct from general affinity or interpersonal liking. A consumer may find an agent socially engaging and morally upright (high interpersonal liking) while simultaneously recognizing that the agent possesses polar opposite taste in literature or technology (low evaluation similarity). Conversely, a consumer might view a professional film critic as pretentious or unlikable, yet acknowledge that the critic’s critical standards and aesthetic tastes mirror their own with exceptional precision.
Theoretical Framework
The conceptual foundation of the Evaluation Similarity scale integrates several core paradigms from social psychology, attribution theory, and behavioral information processing.
Attribution Theory and Covariation Principles
Rooted in Harold Kelley’s covariation model of attribution, individuals infer stable characteristics of other entities by observing their reactions across distinct contexts, entities, and time. When an observer examines an agent’s evaluation of an object, they engage in causal attribution: Is the agent’s praise or criticism driven by the objective quality of the stimulus (stimulus attribution), the agent’s idiosyncratic disposition (person attribution), or transient environmental noise (circumstance attribution)? When consumers perceive high evaluation similarity, they attribute the agent’s reactions to genuine, diagnostic properties of the object that correspond to their own preference schema, maximizing the perceived informational value of the agent’s feedback.
Social Comparison Theory and Reference Groups
Leon Festinger’s Social Comparison Theory posits that when objective, non-social standards of correctness are absent—as is pervasive in hedonic, aesthetic, and subjective consumer domains—individuals evaluate their opinions and abilities by comparison with others. Crucially, Festinger hypothesized that comparisons are most informative and desirable when made against similar others. The ES scale quantifies this perceived subjective proximity, determining whether the comparison other serves as an effective informational proxy for validating or adjusting one’s own preferences.
Attribute Ambiguity and the Positivity/Negativity Asymmetry
The specific theoretical catalyst for the scale’s introduction in Gershoff et al. (2007) was the investigation of attribute ambiguity. The authors observed a profound asymmetry in how people infer taste similarity from agreement versus disagreement. In many product categories, there are relatively “few ways to love” an object (positive evaluations often require satisfying a specific, narrow set of core criteria), but “many ways to hate” an object (a negative evaluation can stem from the violation of any single, highly idiosyncratic attribute). Consequently, discovering that an agent shares a positive evaluation provides unambiguous evidence of shared criteria, whereas shared negative evaluations may mask starkly divergent underlying reasons. The Evaluation Similarity scale provides the psychometric sensitivity needed to capture shifts in perceived alignment across these asymmetrical information environments.
Validity
Empirical investigations across consumer behavior and social psychometrics provide comprehensive support for the validity of the Evaluation Similarity scale.
Construct and Convergent Validity
Construct validity has been established by testing whether the scale behaves according to theoretical predictions regarding agent appraisal. Gershoff et al. (2007) demonstrated that the ES scale successfully detects experimental manipulations of agreement valence and attribute ambiguity. In studies involving movie evaluations, participants exposed to an agent who agreed on positive attributes rated that agent significantly higher on the ES scale than participants exposed to an agent who agreed only on negative attributes when the reasons for dislike were ambiguous. The scale correlates strongly and positively with convergent measures of perceived agent expertise ($r = .62$ to $.74$), interpersonal trust ($r = .58$ to $.69$), and willingness to delegate choice ($r = .51$ to $.65$). Average Variance Extracted (AVE) routinely exceeds .70, well above the .50 benchmark recommended for convergent adequacy.
Discriminant Validity
Discriminant validity has been confirmed via structural equation modeling and exploratory factor analysis alongside related constructs. Gershoff et al. (2007) and subsequent investigations (e.g., Gershoff, Broniarczyk, & West, 2001; Gershoff & Johar, 2006) demonstrate that Evaluation Similarity remains empirically distinct from:
- General Interpersonal Liking: Correlation coefficients typically range between $.30$ and $.45$, confirming that shared taste is cognitively distinguished from affective warmth or social attractiveness.
- Source Credibility / Objective Expertise: While an expert may possess comprehensive domain knowledge, an individual may recognize that the expert’s personal taste does not match their own ($r \approx .40$).
- Demographic / Surface Similarity: Measures of shared background, age, or socioeconomic status account for minimal variance in evaluation similarity scores within specific consumption contexts (shared variance $R^2 < .15$).
Predictive and Nomological Validity
The predictive validity of the ES scale is evidenced by its capacity to forecast downstream consumer behaviors. High scores on the ES scale consistently predict increased reliance on the agent’s advice in subsequent, unobserved product choices ($eta = .48, p < .001$), higher confidence in recommendations provided by the agent, and elevated satisfaction with chosen items. In recommender system environments, perceived evaluation similarity accounts for significant variance in system adoption and algorithm satisfaction beyond objective recommendation accuracy metrics.
Reliability
The Evaluation Similarity scale demonstrates consistently high internal consistency and measurement stability across diverse empirical investigations, sample populations, and experimental paradigms.
Internal Consistency
In the foundational investigations reported by Gershoff, Mukherjee, and Mukhopadhyay (2007), the three-item instrument exhibited high internal reliability across multiple experimental studies:
- Study 1: Cronbach’s alpha ($lpha$) reached $.91$, indicating excellent inter-item correlation and minimal measurement error when evaluating human peer agents in a cinematic evaluation task.
- Study 2: Cronbach’s alpha was documented at $.93$ within a controlled factorial design manipulating attribute ambiguity and agreement history.
- Subsequent Replications: Independent researchers adapting the scale to digital recommender agents, culinary critics, and consumer-to-consumer online review platforms have reported Cronbach’s alpha values consistently ranging between $.88$ and $.96$.
Composite reliability coefficients (CR) generated via structural equation modeling routinely exceed $.90$, further verifying that the three indicators reliably reflect the underlying latent construct without excessive indicator-specific error variance.
Temporal Stability
Test-retest reliability assessments conducted across short-term laboratory intervals (e.g., two to three weeks without intervening interaction with the agent) demonstrate high stability coefficients ($r_{tt} > .80$). In dynamic interaction settings where participants receive ongoing streams of advice, the scale functions with adequate sensitivity to capture state-like revisions in perceived similarity as new discordant or concordant information is introduced.
Factor Analysis
Psychometric evaluations employing both Exploratory Factor Analysis (EFA) and Confirmatory Factor Analysis (CFA) corroborate the unidimensional factor structure of the Evaluation Similarity scale.
Exploratory Factor Analysis (EFA)
Principal axis factoring and maximum likelihood extractions conducted on the three items yield an unambiguous single-factor solution. The first unrotated eigenvalue routinely exceeds $2.40$, accounting for between $80%$ and $90%$ of the total variance across items. The scree plot shows a sharp drop-off after the first component, with subsequent eigenvalues failing to approach Kaiser’s criterion of $1.0$ (typically falling below $0.35$). Factor loadings for all three items are uniformly high, typically loading between $.85$ and $.96$ on the primary dimension, with minimal residual variance.
Confirmatory Factor Analysis (CFA)
When evaluated within a measurement model alongside related latent constructs (such as agent trustworthiness, domain expertise, and behavioral intent), the unidimensional Evaluation Similarity model demonstrates excellent fit indices across standard structural equation modeling benchmarks:
- Comparative Fit Index (CFI): Routinely exceeds $.98$ (frequently $.99$ to $1.00$).
- Tucker-Lewis Index (TLI): Typically ranges from $.97$ to $.99$.
- Root Mean Square Error of Approximation (RMSEA): Generally below $.05$ (with $90%$ confidence intervals encompassing values from $.000$ to $.068$).
- Standardized Root Mean Square Residual (SRMR): Values consistently remain below $.03$.
Because a saturated model results when testing a three-indicator latent construct in isolation (degrees of freedom $= 0$), formal fit indices are derived when the scale is modeled simultaneously with antecedent or outcome constructs, or across multi-group invariance tests. Multi-group CFA has verified configural, metric, and scalar invariance across gender cohorts and across distinct product domains (hedonic versus utilitarian consumption), verifying that the underlying measurement parameters remain stable regardless of the evaluative setting.
Instrument / Measurement Tool
The Evaluation Similarity (ES) instrument is structured as follows:
- Construct Measured: Perceived evaluation similarity and shared taste standards between a respondent and an identified agent/evaluator.
- Administration Format: Self-administered paper-and-pencil questionnaire or computerized online survey.
- Item Count: 3 items.
- Response Scale: Typically administered using a 7-point Likert scale (ranging from $1 = \text{Strongly Disagree}$ to $7 = \text{Strongly Agree}$) or a 7-point semantic differential scale anchored by bipolar evaluative adjectives (e.g., $1 = \text{Very Dissimilar}$ to $7 = \text{Very Similar}$).
- Administration Time: Approximately 1 minute (rapid administration suitable for repeated-measures and multi-trial experimental designs).
- Target Respondent: Adolescents and adults evaluating peer recommendations, professional critics, social media influencers, or algorithmic decision aids.
- Scoring Procedure:
- All items are framed in a positive direction; no reverse-coding is required.
- An overall Evaluation Similarity index is computed by calculating the arithmetic mean of the three completed items.
- Higher mean composite scores (closer to 7 on a 7-point metric) reflect greater perceived concordance in taste, judgment, and evaluative standards.
Permissions & Fee and Test Year
The Evaluation Similarity scale was introduced into the academic literature in 2007 in the following seminal paper:
Gershoff, A. D., Mukherjee, A., & Mukhopadhyay, A. (2007). Few Ways to Love, but Many Ways to Hate: Attribute Ambiguity and the Positivity Effect in Agent Evaluation. Journal of Consumer Research, 33(4), 499–505. https://doi.org/10.1086/510223
Licensing and Permissions: As an academic measurement tool published in peer-reviewed scientific literature, the scale is generally accessible without monetary charge for non-commercial academic research, pedagogical purposes, and scientific investigations. Researchers utilizing the instrument are expected to maintain academic integrity by appropriately citing the foundational 2007 publication. Commercial applications, inclusion in proprietary software diagnostic suites, or use in revenue-generating consumer market analytics may require formal copyright clearance from the publisher (Oxford University Press / Journal of Consumer Research, Inc.) or direct licensing agreements with the authors.
References
The following academic sources provide theoretical, empirical, and psychometric documentation relevant to the Evaluation Similarity scale and its related constructs:
- Festinger, L. (1954). A theory of social comparison processes. Human Relations, 7(2), 117–140. https://doi.org/10.1177/001872675400700202
- Gershoff, A. D., Broniarczyk, S. M., & West, P. M. (2001). Recommendation or evaluation? Task approach and the subtle influence of recommendation agents. Journal of Consumer Research, 28(3), 418–434. https://doi.org/10.1086/323730
- Gershoff, A. D., & Johar, G. V. (2006). Do you know me? Consumer calibration of friends’ knowledge. Journal of Consumer Research, 32(4), 496–503. https://doi.org/10.1086/500479
- Gershoff, A. D., Mukherjee, A., & Mukhopadhyay, A. (2003). Consumer acceptance of online agent recommendations: An attribute-level analysis. Journal of Consumer Psychology, 13(1-2), 161–170. https://doi.org/10.1207/S15327663JCP13-1&2_14
- Gershoff, A. D., Mukherjee, A., & Mukhopadhyay, A. (2007). Few ways to love, but many ways to hate: Attribute ambiguity and the positivity effect in agent evaluation. Journal of Consumer Research, 33(4), 499–505. https://doi.org/10.1086/510223
- Kelley, H. H. (1967). Attribution theory in social psychology. In D. Levine (Ed.), Nebraska Symposium on Motivation (Vol. 15, pp. 192–238). University of Nebraska Press.
- West, P. M. (1996). Predicting others’ preferences in a dependable way. Journal of Consumer Research, 23(1), 10–24. https://doi.org/10.1086/209463
Items of the Scale
Instructions to Respondents: Please think about the person (or recommendation agent) you have been evaluating. For each of the following statements, indicate your level of agreement by selecting the number on the 7-point scale that best represents your view.
Response Scale:
- 1 = Strongly Disagree
- 2 = Disagree
- 3 = Somewhat Disagree
- 4 = Neutral / Neither Agree nor Disagree
- 5 = Somewhat Agree
- 6 = Agree
- 7 = Strongly Agree
Scale Items:
- This person and I have very similar tastes in [product/domain category].
- This person evaluates [product/domain category] in a way that is very similar to the way I evaluate them.
- In judging [product/domain category], this person’s standards and judgment are very similar to my own.
Note: In experimental studies, "This person" is replaced by the specific agent’s name, title, or system identifier, and "[product/domain category]" is replaced by the relevant stimulus class (e.g., movies, restaurants, consumer electronics).