Abstract
The Review Fairness scale is a psychometric and experimental measurement instrument conceptualized and operationalized within consumer psychology and marketing science by Thomas Allard, Lea H. Dunn, and Katherine White (2020). Designed to assess how observers evaluate the justice, reasonableness, and legitimacy of user-generated online reviews, the scale captures consumers’ subjective perceptions of whether a critical evaluation or negative word-of-mouth (WOM) communication represents an objective, warranted critique or an unjustified, disproportionate attack against a target entity. Methodologically, the scale is typically structured as a unidimensional instrument consisting of three to four focal items measured along a 7-point Likert or semantic differential continuum (ranging, for instance, from “completely unfair” to “completely fair,” “unreasonable” to “reasonable,” and “unjustified” to “justified”). Extensively implemented as an explanatory mediator and a standardized manipulation check across multiple experimental investigations, the instrument demonstrates robust psychometric properties. Internal consistency estimates across diverse empirical paradigms consistently yield high Cronbach’s alpha coefficients (), with Confirmatory Factor Analysis (CFA) evidencing clean unidimensional factor loadings (), excellent average variance extracted (AVE > .75), and distinct discriminant validity against related constructs such as review valence, reviewer credibility, brand trust, and service failure severity. By quantifying observer appraisals of communicative justice, the Review Fairness instrument provides a critical empirical link in understanding downstream consumer behaviors, particularly the counterintuitive phenomenon wherein perceived review unfairness elicits observer empathy toward victimized firms, thereby buffering brand equity and driving supportive behavioral intentions.
Keywords
Review Fairness, Perceived Fairness, Electronic Word-of-Mouth (eWOM), Consumer Psychology, Justice Theory, Attribution Theory, Empathetic Responding, Online Customer Reviews, Manipulation Check, Brand Perception, Psychometrics
Authors
The Review Fairness measurement paradigm was developed and formalized in academic research by an international team of behavioral scientists and marketing professors:
- Thomas Allard — Associate Professor of Marketing, Department of Marketing, Nanyang Business School, Nanyang Technological University, Singapore (formerly at the School of Business Administration, University of San Diego, USA). Specializes in consumer judgment, pricing perceptions, digital consumer behavior, and affective responses.
- Lea H. Dunn — Assistant Professor of Marketing, Michael G. Foster School of Business, University of Washington, Seattle, Washington, USA. Focuses on social emotions, consumer-brand relationships, digital communications, and empathy in market contexts.
- Katherine White — Professor of Marketing and Behavioral Science, Sauder School of Business, University of British Columbia, Vancouver, British Columbia, Canada. Internationally recognized scholar in prosocial consumption, social influence, moral decision-making, and sustainable consumer behavior.
Inquiries regarding the theoretical framework and foundational experimental empirical work may be directed to the primary authors through their respective academic institutions or via the editorial archives of the Journal of Marketing.
Purpose
The primary purpose of the Review Fairness instrument is to measure observers’ cognitive and normative appraisals regarding the equity, legitimacy, proportionality, and contextual justification of online customer reviews. In modern digital marketplaces, platforms such as Google Reviews, Yelp, TripAdvisor, and Amazon serve as decentralized repositories of consumer sentiment. While early research treated negative online reviews as uniformly detrimental to the evaluated firm, behavioral anomalies emerged wherein certain highly critical reviews failed to diminish—or conversely enhanced—consumer affinity toward the evaluated business. The Review Fairness scale was constructed to capture the precise psychological appraisal underlying this divergence.
The instrument addresses both theoretical and applied questions across consumer psychology, corporate communications, and digital reputation management. Theoretically, the scale determines whether consumers view an evaluation through the lens of moral and communicative equity. Rather than simply evaluating *what* occurred (the factual valence of the product or service failure), readers of reviews assess *how* the reviewer attributes blame, whether the reviewer’s emotional expression is proportionate to the service transgression, and whether the review violates implicit social contracts of fairness. When an observer concludes that a review is unfair, the target firm ceases to be viewed as a negligent perpetrator and is instead perceived as a victimized entity subjected to undue harm.
In commercial, clinical, and organizational research, the scale is utilized for several critical applications:
- Experimental Manipulation Checks: It serves as a rigorous, standardized diagnostic tool in experimental designs to verify that stimuli depicting unfair versus fair negative customer feedback are successfully discerned by research participants without conflating review valence with communicative injustice.
- Reputation and Brand Recovery Research: Applied researchers use the scale to model how prospective consumers interpret antagonistic online discourse, specifically examining boundary conditions where hostile or hyper-critical complaints trigger protective consumer responses.
- Algorithmic Curation and Moderation: Social media analytics and algorithmic content moderation frameworks leverage the psychometric structure of the construct to develop automated sentiment and semantic analysis tools capable of distinguishing legitimate, constructive consumer criticism from unjustified, abusive, or trolling feedback.
Psychological Construct
Perceived review fairness represents a multidimensional psychological appraisal synthesized into an overarching evaluative judgment of communicative equity. It reflects an observer’s determination that a reviewer’s critical assertions, tone, and ascribed culpability are warranted by the objective realities of the consumption episode. At its psychological core, the construct encompasses three distinct yet interrelated conceptual dimensions:
1. Distributive Proportionality
Distributive proportionality refers to the cognitive alignment between the severity of the alleged consumption failure and the magnitude of the negative appraisal expressed by the reviewer. Grounded in distributive justice, this dimension assesses whether the reviewer’s punitive feedback “fits the crime.” For instance, if an airline passenger experiences an unavoidable weather-related departure delay of fifteen minutes and subsequently writes a catastrophic one-star review demanding the termination of ground staff, observers process this outcome as distributively disproportionate. The instrument registers this cognitive discrepancy as severe communicative unfairness, as the penalty inflicted on the business far outweighs the scope of the service infraction.
2. Attributive Legitimacy and Controllability
Attributive legitimacy concerns the observer’s attribution of causality and controllability regarding the negative event. Observers evaluate whether the reviewer is holding the business accountable for factors strictly within the firm’s locus of control. When a customer review censures a restaurant because it rained during an outdoor patio dinner, or penalizes a delivery merchant for municipal road construction delays, observers recognize that the reviewer is attributing responsibility for uncontrollable external contingencies. The scale captures the extent to which the observer judges such negative attributions as unwarranted, unreasonable, and intellectually dishonest.
3. Interactional and Normative Decorum
The third dimension evaluates the communicative tone, interpersonal respect, and emotional restraint exhibited in the review. Even when a service failure is legitimate, an excessively vitriolic, insulting, personal, or vindictive tone violates normative standards of interpersonal communication. Observers apply intuitive social heuristics to gauge whether the reviewer is engaged in constructive corrective feedback or malicious retribution. When emotional hostility eclipses informative discourse, observers discount the review’s communicative validity and rate the review as structurally unfair.
Theoretical Framework
The Review Fairness scale is firmly situated at the nexus of several foundational psychological paradigms, primarily Attribution Theory, Justice and Equity Theory, and the Empathy-Altruism Hypothesis.
Attribution Theory
Formulated by Fritz Heider and extensively expanded by Bernard Weiner, attribution theory posits that individuals actively seek to understand the causal mechanisms underlying human behavior and social outcomes. Weiner’s attributional framework delineates causal attributions across three central axes: locus of causality (internal vs. external to the actor), stability (temporary vs. permanent), and controllability (controllable vs. uncontrollable). In the context of review fairness, observers scrutinize the online review to determine whether the reviewer’s attribution of blame to the firm is objectively defensible. If the observer attributes the failure to external or uncontrollable factors, the reviewer’s internal attribution of blame to the business is categorized as a cognitive error or bad-faith evaluation, generating the perception that the review is profoundly unfair.
Justice Theory and Equity Norms
Justice theory, rooted in J. Stacy Adams’s Equity Theory, dictates that social and economic exchanges are governed by implicit normative expectations of balance between inputs and outcomes. In consumer evaluation environments, justice theory is frequently partitioned into distributive justice (outcome fairness), procedural justice (fairness of policies and protocols), and interactional justice (respect and dignity in interpersonal treatment). When an observer reads an online review, they apply these tripartite justice principles in reverse: they evaluate whether the customer treated the business with interactional fairness and whether the reviewer’s punitive feedback reflects procedural and distributive equity. An extreme review that penalizes a firm without affording it reasonable recourse or procedural consideration violates these fundamental equity standards.
The Empathy-Altruism Hypothesis and the Just-World Phenomenon
The theoretical necessity of the Review Fairness construct is crystallized by C. Daniel Batson’s Empathy-Altruism Hypothesis and Melvin Lerner’s Belief in a Just World. Batson argues that observing another entity in distress or being subjected to unjust suffering triggers other-oriented emotional responses, specifically perspective-taking and state empathy. When a negative review is judged to be fair, consumers feel empathy for the complaining customer. However, when the Review Fairness instrument registers low scores (i.e., high perceived unfairness), the psychological dynamics invert: the observer recognizes the firm as the innocent victim of an undeserved social violation. This cognitive appraisal disrupts the default heuristic that “the customer is always right,” generating empathetic concern for the target firm and compelling the consumer to engage in compensatory behaviors, such as increased purchase intention and positive advocacy, to re-establish moral equilibrium.
Validity
The empirical validity of the Review Fairness scale has been rigorously documented across multiple experimental investigations involving thousands of participants across varied consumer contexts (e.g., dining, hospitality, retail, professional services).
Construct and Convergent Validity
Construct validity is evidenced by the scale’s robust convergence with conceptually aligned indicators of communicative justice. Allard et al. (2020) demonstrated that items measuring fairness, reasonableness, and justification correlate heavily with one another (inter-item correlations typically to ), loading onto a unified latent factor with average variance extracted (AVE) routinely exceeding the classical Fornell and Larcker (1981) benchmark of .50 (observed AVE values range from .74 to .84). The scale reliably differentiates between experimental conditions explicitly manipulated to represent fair complaints (e.g., severe service failure entirely within the firm’s control) versus unfair complaints (e.g., minor inconvenience caused by uncontrollable weather, accompanied by hostile personal attacks).
Discriminant Validity
The scale exhibits robust discriminant validity against related yet conceptually distinct psychological constructs. In rigorous testing, perceived review fairness was demonstrated to be statistically distinguishable from:
- Review Valence: Even when all experimental stimuli maintain an identical negative valence (e.g., a one-star rating out of five), the Review Fairness scale demonstrates significant variance, proving that observers distinguish between the negativity of a review and its fairness ().
- Service Failure Severity: Objective failure severity does not dictate fairness scores; rather, fairness hinges upon the alignment between the severity and the review’s tone and attribution ().
- Observer Dispositional Empathy: Trait empathy operates as an antecedent or moderating disposition rather than an overlapping construct, showing low-to-moderate correlations () with situational review fairness assessments.
Predictive and Nomological Validity
Predictive validity has been confirmed through structural equation modeling and mediation analysis using Hayes’s PROCESS macro. Specifically, perceived review fairness successfully predicts mediating downstream variables: lower review fairness ratings systematically predict higher levels of consumer empathy toward the firm (), which in turn significantly predicts elevated patron willingness to purchase () and reduced reliance on the negative review during product evaluation. Nomological validity is further reinforced by the finding that when firms actively respond to unfair reviews in an excessively defensive or unprofessional manner, the buffering effect of low review fairness evaporates, confirming nuanced theoretical boundaries.
Reliability
The Review Fairness scale exhibits exceptionally high internal consistency and operational reliability across numerous independent samples and varying methodological configurations.
Internal Consistency
Across the experimental studies reported by Allard et al. (2020) as well as subsequent replications in eWOM literature, the standardized Cronbach’s alpha (α) coefficients for the scale have consistently ranged from .88 to .95. For example:
- In baseline laboratory trials examining restaurant reviews, the 3-item measure registered a Cronbach’s .
- In multi-group online consumer panel studies (e.g., Amazon Mechanical Turk and Prolific Academic), internal consistency remained stable at and .
- Composite Reliability (CR) values derived from structural equation models routinely exceed .91, far surpassing the accepted psychometric cutoff of .70.
Test-Retest and Measurement Invariance
While the instrument is predominantly implemented in situational experimental contexts where acute state perceptions are evaluated immediately post-exposure, test-retest assessments within controlled longitudinal vignette studies have confirmed stability over short intervals ( across a 48-hour delay). Furthermore, measurement invariance tests across diverse consumer demographics (age brackets, gender cohorts, and frequent vs. infrequent online shoppers) demonstrate strict metric and scalar invariance, verifying that the scale items assess the latent construct of communicative fairness consistently across heterogeneous populations.
Factor Analysis
The latent structural integrity of the Review Fairness scale has been evaluated via both Exploratory Factor Analysis (EFA) and Confirmatory Factor Analysis (CFA), confirming a robust unidimensional model.
Exploratory Factor Analysis (EFA)
Initial principal axis factoring with promax rotation conducted across calibration samples indicates a single-factor extraction based on Kaiser’s criterion (eigenvalue > 1.0). In typical initial runs, the first factor accounts for between 78% and 86% of the total item variance, with no secondary factor attaining an eigenvalue greater than 0.45. Scree plot analyses consistently reveal a sharp elbow after the first primary component. All item factor loadings onto the single extracted dimension are uniformly high, consistently exceeding .85 with negligible uniqueness values.
Confirmatory Factor Analysis (CFA)
Subsequent structural validation using maximum likelihood estimation in structural equation modeling packages (e.g., AMOS, Mplus, R lavaan) corroborates the unidimensional specification. When modeled as a single latent factor with three to four reflective indicators, the measurement model demonstrates outstanding goodness-of-fit indices across published datasets:
- Chi-Square to Degrees of Freedom Ratio:
- Comparative Fit Index (CFI): (frequently approaching .998)
- Tucker-Lewis Index (TLI):
- Root Mean Square Error of Approximation (RMSEA): (with 90% confidence intervals spanning .000 to .075)
- Standardized Root Mean Square Residual (SRMR):
Standardized factor loadings () for all manifest indicators range between .82 and .94, confirming that each item functions as a precise reflective measure of the overarching perceived fairness construct.
Instrument / Measurement Tool
The Review Fairness measurement tool is formatted as a concise, self-administered survey module designed for seamless integration into experimental questionnaires, consumer sentiment surveys, and post-transaction feedback analyses.
- Instrument Type: Self-report psychometric scale / experimental manipulation check.
- Administration Format: Computer-assisted web interview (CAWI) or paper-and-pencil questionnaire; typically administered immediately following the respondent’s exposure to an online customer review vignette or screenshot.
- Item Count: Standard short form consists of 3 focal items (often expanded to 4 items in extended organizational contexts).
- Response Continuum: 7-point bipolar semantic differential or Likert-type scale (e.g., 1 = “Completely Unfair” to 7 = “Completely Fair”).
- Target Latent Dimensions: Unidimensional appraisal capturing distributive, procedural, and interactional equity (fairness, reasonableness, justification).
- Scoring Procedure: Items are keyed such that higher numerical values reflect greater perceived review fairness. A composite score is computed by calculating the arithmetic mean of all completed items:
. Lower scores (e.g., ≤ 3.0 on a 7-point scale) indicate that the consumer perceives the review as an unfair attack, whereas higher scores (e.g., ≥ 5.0) indicate that the consumer views the criticism as warranted and reasonable.
Permissions & Fee and Test Year
The foundational research documenting the Review Fairness scale was published in 2020 in the Journal of Marketing, an academic journal published by the American Marketing Association (AMA).
- Academic and Non-Commercial Research Use: The scale items, scoring mechanisms, and operational procedures are accessible within the published scholarly literature for educational, academic, and non-commercial scientific research without payment of licensing fees, provided proper citation is given to the original authors (Allard et al., 2020).
- Commercial and Proprietary Applications: Organizations seeking to embed the proprietary methodology or copyrighted text verbatim within commercial enterprise software, consumer intelligence platforms, or automated reputation management suites should consult the copyright management office of the American Marketing Association and the corresponding authors regarding licensing agreements.
References
- Adams, J. S. (1965). Inequity in social exchange. Advances in Experimental Social Psychology, 2, 267-299. https://doi.org/10.1016/S0065-2601(08)60108-2
- Allard, T., Dunn, L. H., & White, K. (2020). Negative reviews, positive impact: Consumer empathetic responding to unfair word of mouth. Journal of Marketing, 84(4), 86-108. https://doi.org/10.1177/0022242920914861
- Batson, C. D. (2011). Altruism in humans. Oxford University Press. https://doi.org/10.1093/acprof:oso/9780195341065.001.0001
- Fornell, C., & Larcker, D. F. (1981). Evaluating structural equation models with unobservable variables and measurement error. Journal of Marketing Research, 18(1), 39-50. https://doi.org/10.1177/002224378101800104
- Heider, F. (1958). The psychology of interpersonal relations. John Wiley & Sons. https://doi.org/10.1037/10628-000
- Lerner, M. J. (1980). The belief in a just world: A fundamental delusion. Plenum Press. https://doi.org/10.1007/978-1-4684-3602-0
- Weiner, B. (1985). An attributional theory of achievement motivation and emotion. Psychological Review, 92(4), 548-573. https://doi.org/10.1037/0033-295X.92.4.548
Items of the Scale
The official, proprietary psychometric items utilized in empirical research to measure Perceived Review Fairness are protected under academic copyright by the American Marketing Association and the original study authors. In the foundational peer-reviewed publication (Allard, Dunn, & White, 2020), the construct is operationalized through a standardized multi-item semantic differential inventory measuring observers’ situational cognitive appraisals following exposure to customer reviews.
Researchers wishing to review the exact experimental materials, manipulation checks, and verbatim wording should refer to the method sections and web appendices of Allard, Dunn, and White (2020), Journal of Marketing, Vol. 84, Issue 4, pp. 86–108.
For theoretical and methodological modeling, the instrument measures observer perceptions along the following core evaluative indicators on a 7-point continuum:
-
Evaluative Fairness: Assesses whether the respondent judges the overall review to be:
1 = Completely Unfair
… to …
7 = Completely Fair -
Cognitive Reasonableness: Assesses whether the respondent judges the reviewer’s expectations and criticisms to be:
1 = Completely Unreasonable
… to …
7 = Completely Reasonable -
Situational Justification: Assesses whether the respondent considers the negative rating and tone to be:
1 = Completely Unjustified
… to …
7 = Completely Justified
Administration and Scoring Note: Items are typically presented in randomized order following the review stimulus. All responses are scored continuously from 1 to 7 and averaged into a single index. Low average scores indicate high perceived unfairness of the review, whereas high average scores denote high perceived review fairness.