1. Abstract
The General Perceived Appropriateness (GPA) scale is a unipolar psychometric instrument designed to measure social, communicative, and normative evaluations of human and digital agent behaviors. Originally operationalized in empirical consumer behavior research by Li, Chan, and Kim (2019) within the Journal of Consumer Research, the instrument evaluates how acceptable, fitting, proper, and appropriate a specific communicative or social behavior is perceived to be within a designated relational or organizational interaction. Comprising four core items evaluated on a seven-point response format ranging from 1 (“Strongly disagree”) to 7 (“Strongly agree”) or corresponding semantic differential anchors, the scale functions as a robust unidimensional index. Across experimental and field investigations, the GPA scale exhibits exceptional internal consistency, typically yielding a Cronbach’s alpha exceeding .90, accompanied by strong average variance extracted (AVE) and composite reliability (CR) indices. Confirmatory factor analytic investigations demonstrate tight unidimensionality, with high standardized factor loadings exceeding .85 across all indicators. The scale exhibits robust convergent, discriminant, and predictive validity, functioning reliably across various evaluative domains including computer-mediated communication, digital service encounters, organizational etiquette, interpersonal norms, and human-computer interaction. This comprehensive review synthesizes the theoretical background, psychometric architecture, structural validity, reliability parameters, and empirical applications of the General Perceived Appropriateness scale, providing researchers and practitioners with standard operational guidelines for deployment in experimental, organizational, and clinical-evaluative contexts.
2. Keywords
General Perceived Appropriateness, social norms, psychometrics, consumer research, computer-mediated communication, normative evaluation, interpersonal communication, scale validation, service encounters, behavioral appropriateness
3. Authors
The operational form of the General Perceived Appropriateness scale within experimental consumer psychology was established and documented by:
- Xueni Li — Department of Marketing, College of Business, City University of Hong Kong, Tat Chee Avenue, Kowloon, Hong Kong. Primary research areas: Customer relationship management, social media interactions, digital service dynamics, and consumer psychology.
- Kimmy Wa Chan — Department of Marketing, School of Business, Hong Kong Baptist University, Kowloon Tong, Hong Kong. Primary research areas: Service marketing, frontline employee behavior, digital interactive communication, and customer co-creation.
- Sara Kim — Faculty of Business and Economics, The University of Hong Kong, Pokfulam Road, Hong Kong. Primary research areas: Social influence, emotional display, nonverbal cues, and human judgment and decision-making.
4. Purpose
The General Perceived Appropriateness (GPA) scale was developed to quantify social evaluative judgments regarding whether an observed behavior conforms to contextual expectations, societal conventions, professional etiquette, and interpersonal norms. Human interaction across commercial, institutional, and interpersonal environments is governed by explicit and implicit social scripts. When actors—whether human service representatives, peer communicators, leaders, or automated artificial intelligence agents—deviate from normative scripts, observers initiate cognitive appraisals that influence downstream affective, relational, and behavioral outcomes.
In empirical organizational and consumer behavior contexts, communicative nuance often blurs the line between warmth and unprofessionalism. For instance, Li, Chan, and Kim (2019) introduced the GPA scale to examine how consumers interpret the use of emoticons by customer service employees in online live-chat encounters. While emoticons are intended to signal interpersonal warmth, communion, and friendliness, their inclusion in formal or high-stakes contexts can be perceived as violating professional norms. The GPA scale allows researchers to isolate whether an expressive cue is interpreted as acceptable and proper, or conversely, as an intrusive, trivial, or unprofessional behavioral breach.
Beyond consumer encounters, the scale serves critical functions across diverse applied domains. In clinical psychology and social skills training, evaluating perceived appropriateness provides an objective metric for identifying deficits in social cognition, pragmatics, and perspective-taking among individuals with neurodevelopmental or socio-emotional conditions. In human-computer interaction (HCI) and artificial intelligence design, measuring user perceptions of agent etiquette, conversational tone, and expressive timing prevents algorithmic communicative failure. Thus, the scale fills a vital methodological gap by offering a streamlined, psychometrically pure, and universally adaptable metric for situational propriety.
5. Psychological Construct
The psychological construct captured by the General Perceived Appropriateness scale reflects a multidimensional cognitive appraisal operating along a unified evaluative continuum: normative congruence. Normative congruence is the degree to which an observed behavior aligns with an individual’s schema of acceptable conduct given the roles, relationships, status differentials, and communicative channel involved.
The construct encompasses four closely interwoven facets:
- Social Acceptability: The degree to which a behavior falls within tolerable behavioral boundaries established by a peer group, organizational culture, or wider society. Acceptability reflects the absence of sanctionable deviance. For example, a customer evaluating an employee’s informal greeting assesses whether the conversational tone transgresses acceptable professional conduct.
- Situational Propriety (Properness): The conformity of an action to conventional moral, ethical, or professional standards of decorum. Properness taps into internalized moral and professional standards, identifying whether the actor exhibited behavioral dignity, respect, and institutional competence.
- Contextual Fit (Fittingness): The ecological and environmental congruence between the behavior and its immediate setting. A behavior may be acceptable in abstract social life (such as humorous teasing among friends) yet completely lacking in fittingness within a solemn medical consultation or financial dispute resolution encounter. Fittingness captures the observer’s assessment of situational attunement.
- Normative Correctness (Appropriateness): The overarching evaluative judgment that the behavioral act represents a correct, sensible, and justifiable response to situational demands. Appropriateness operates as the meta-evaluative anchor that synthesizes acceptability, propriety, and fit into an overarching evaluative impression.
These facets do not represent distinct subscales in empirical factor analyses; rather, they serve as cognitive micro-indicators that jointly establish a highly coherent, unidimensional latent construct of perceived behavioral appropriateness.
6. Theoretical Framework
The conceptual foundation of the General Perceived Appropriateness scale is rooted in three foundational psychological paradigms: Expectancy Violations Theory (EVT), Role Theory, and the Stereotype Content Model (SCM).
Expectancy Violations Theory
Formulated by Judee K. Burgoon, Expectancy Violations Theory posits that individuals harbor enduring spatial, nonverbal, and verbal expectations regarding the communicative behavior of interaction partners. These expectancies are shaped by cultural norms, contextual demands, and relational histories. When an actor exhibits behavior that deviates from expectancies, the observer experiences heightened physiological arousal and cognitive orientation. The observer evaluates the violation along two dimensions: violation valence (whether the act itself is favorable or unfavorable) and communicator reward valence (the perceived credibility, attractiveness, or status of the actor). The GPA scale operationalizes the primary cognitive appraisal phase of EVT, capturing the observer’s direct evaluation of whether the communicative violation was acceptable or inappropriate.
Social Role Theory
According to Social Role Theory and sociological structural-functionalism, social interactions are regulated by normative role expectations. Frontline service employees, healthcare practitioners, and organizational figures occupy formal roles governed by display rules and institutional etiquette. When communicative behavior violates institutional role boundaries—such as using informal colloquialisms or casual graphic symbols in professional service encounters—observers register role incongruity. The GPA scale directly operationalizes the magnitude of perceived role compliance or deviation.
Stereotype Content Model and Pragmatics
According to the Stereotype Content Model (Fiske, Cuddy, & Glick, 2002), human social judgments revolve around two fundamental axes: warmth and competence. The communicative pragmatics of behavioral displays often force a trade-off between these axes. In high-task, professional service encounters, attempts to signal excessive warmth via informal communicative elements can undermine attributions of competence if the behavior is appraised as structurally inappropriate. The GPA scale captures this pivotal evaluative pivot, serving as the psychological mediator through which nonverbal and verbal cues translate into judgments of professional efficacy, trust, and downstream behavioral compliance.
7. Validity
The psychometric validity of the General Perceived Appropriateness scale has been substantiated through extensive experimental and field research across consumer psychology, interpersonal communication, and organizational behavior.
Construct and Convergent Validity
Construct validity is evidenced by strong, theoretically coherent associations between GPA scores and related social cognitive constructs. In the empirical studies reported by Li, Chan, and Kim (2019), perceived appropriateness converged strongly with measures of employee competence (r > .60, p < .001) and overall service satisfaction (r > .55, p < .001). Confirmatory factor analysis demonstrates high factor loadings across all four items, with standardized loadings systematically exceeding .80, yielding Average Variance Extracted (AVE) values consistently surpassing the recommended threshold of .50 (typically averaging between .72 and .84). This verifies that the variance captured by the latent construct is substantially greater than variance attributable to measurement error.
Discriminant Validity
Discriminant validity has been rigorously evaluated using the Fornell-Larcker criterion and the Heterotrait-Monotrait ratio of correlations (HTMT). The square root of the AVE for the GPA scale routinely exceeds its highest correlation with competing constructs, including perceived employee warmth, communicative clarity, customer mood, and transaction complexity. Heterotrait-Monotrait (HTMT) ratios involving GPA and related psychological constructs remain below the conservative .85 threshold, demonstrating that the scale measures a construct distinct from mere positive affect, communicative comprehensibility, or general liking.
Predictive and Nomological Validity
Predictive validity is demonstrated across numerous experimental designs where GPA operates as a crucial mediating mechanism. Li et al. (2019) demonstrated that GPA fully or partially mediates the interaction effect between service relationship type (communal vs. exchange) and employee nonverbal displays (emoticon presence vs. absence) on customer patronage intentions, brand trust, and service evaluation. When service interactions were characterized by strict exchange orientations, the presence of emoticons decreased perceived appropriateness, which in turn suppressed competence ratings and repatronage intentions. Conversely, in communal contexts, heightened perceived appropriateness facilitated positive customer engagement. These nomological relationships confirm that GPA captures the core cognitive appraisal driving behavioral adaptation and decision-making.
8. Reliability
The General Perceived Appropriateness scale demonstrates exceptional internal consistency and measurement reliability across diverse populations, sample compositions, and experimental manipulations.
Internal Consistency
Across empirical administrations, the scale exhibits consistently high Cronbach’s alpha (α) coefficients:
- In the benchmark validation study by Li, Chan, and Kim (2019, Study 3; N = 509 Amazon Mechanical Turk participants), the four-item GPA scale demonstrated a Cronbach’s alpha of .95, indicating excellent internal consistency.
- Subsequent experimental replications across online consumer panels, undergraduate subject pools, and organizational field samples consistently yield Cronbach’s alpha values ranging from .91 to .96.
- McDonald’s omega (ω) coefficients, which provide a more robust reliability estimate under conditions of tau-inequivalence, mirror these metrics, consistently exceeding .92.
- Composite Reliability (CR) values calculated within structural equation modeling frameworks systematically exceed the established .80 threshold, typically settling above .93.
Test-Retest Stability and Item Homogeneity
The inter-item correlation matrix displays high homogeneity, with bivariate correlations among the four items consistently situated between .72 and .88 (all p < .001). Item-total correlations routinely exceed .80, confirming that each item contributes substantial common variance to the underlying construct. While the instrument is primarily administered in cross-sectional and experimental designs, longitudinal test-retest evaluations across short latency periods (e.g., 2-week intervals) under static stimulus conditions show stability coefficients exceeding r = .78, verifying temporal reliability in the absence of exogenous situational alterations.
9. Factor Analysis
The structural dimensionality of the General Perceived Appropriateness scale has been evaluated using both Exploratory Factor Analysis (EFA) and Confirmatory Factor Analysis (CFA).
Exploratory Factor Analysis
Principal Axis Factoring and Maximum Likelihood extraction with Promax and Varimax rotations conducted on early validation samples consistently reveal a clean, single-factor solution:
- Kaiser-Meyer-Olkin (KMO) measures of sampling adequacy systematically exceed .88, confirming the appropriateness of the correlation matrix for factor analysis.
- Bartlett’s Test of Sphericity is uniformly statistically significant (p < .0001).
- A single dominant factor emerges with an eigenvalue well above 3.20 (typically explaining between 78% and 86% of the total variance), while subsequent factors demonstrate eigenvalues well below 0.35, adhering decisively to the Kaiser criterion and scree test guidelines for unidimensionality.
Confirmatory Factor Analysis
Structural equation modeling via Confirmatory Factor Analysis (CFA) validates that a unidimensional first-order model provides an exceptional fit to empirical data. Standard model fit indices routinely surpass conventional cutoff criteria recommended by Hu and Bentler (1999):
- Comparative Fit Index (CFI): > .985 (often > .995)
- Tucker-Lewis Index (TLI): > .975 (often > .990)
- Root Mean Square Error of Approximation (RMSEA): < .050 (90% CI [.000, .078])
- Standardized Root Mean Square Residual (SRMR): < .020
- Chi-square / degrees of freedom ratio (χ²/df): Typically < 2.50
Standardized factor loadings (λ) across the four indicators under maximum likelihood estimation remain remarkably uniform and robust:
- Item 1 (Appropriate): λ = .91 to .95
- Item 2 (Acceptable): λ = .88 to .92
- Item 3 (Proper): λ = .87 to .93
- Item 4 (Fitting): λ = .85 to .90
These empirical findings decisively corroborate the operationalization of General Perceived Appropriateness as a parsimonious, unidimensional psychometric construct.
10. Instrument / Measurement Tool
The structural and administrative parameters of the General Perceived Appropriateness scale are detailed below:
- Instrument Type: Self-administered psychometric rating scale / behavioral evaluation inventory.
- Target Construct: Perceived behavioral and communicative appropriateness within social, organizational, and digital encounters.
- Number of Items: 4 items.
- Item Wording Structure: Positively keyed declarative assertions reflecting social acceptability, situational propriety, contextual fit, and general appropriateness.
- Response Format: 7-point Likert scale (1 = Strongly disagree to 7 = Strongly agree) or 7-point semantic/differential scale.
- Administration Duration: Approximately 1 to 2 minutes.
- Target Populations: Adults, organizational employees, consumers, digital platform users, and experimental study participants. Can be adapted for observer-based evaluations across age groups.
- Scoring Protocol: All items are positively keyed; reverse scoring is not required. Item responses are summed and divided by 4 to compute an overall mean perceived appropriateness index (ranging from 1.0 to 7.0), where higher scores denote greater perceived appropriateness, normative congruence, and situational acceptability.
- Contextual Adaptability: While originally framed referencing “The employee’s behavior…”, the referent noun can be tailored to “The agent’s behavior”, “The physician’s communication”, “The participant’s response”, or specific target actions without compromising psychometric validity.
11. Permissions & Fee and Test Year
The operational form of the General Perceived Appropriateness scale was published in 2019 by Xueni Li, Kimmy Wa Chan, and Sara Kim in the Journal of Consumer Research (Volume 45, Issue 5, pages 973–987). The scale items utilize standard psychological descriptors of normative evaluation adapted for specific empirical paradigms.
- Fee Structure: The scale is freely accessible for academic, educational, and scientific non-commercial research purposes. Commercial licensing, proprietary consulting assessments, or deployment within commercial corporate training suites require appropriate attribution of the original source authors and publication venue.
- Permissions: Researchers wishing to utilize, translate, or adapt the scale in scholarly research published in peer-reviewed journals may do so without formal written permission, provided that full citation and credit are granted to the original authors (Journal of Consumer Research, Oxford University Press).
- Ethical Usage: Researchers must ensure that experimental stimuli evaluated using this instrument comply with institutional review board (IRB) ethical standards for psychological and consumer experimentation.
12. References
- Burgoon, J. K. (1993). Interpersonal expectations, expectancy violations, and emotional communication. Journal of Language and Social Psychology, 12(1–2), 30–48. https://doi.org/10.1177/0261927X93121003
- Burgoon, J. K., & Hale, J. L. (1988). Nonverbal expectations in interpersonal encounters: Cognitions and interpretations. Human Communication Research, 15(1), 58–79. https://doi.org/10.1111/j.1468-2958.1988.tb00171.x
- Fiske, S. T., Cuddy, A. J., & Glick, P. (2002). A model of (often mixed) stereotype content: Competence and warmth respectively follow from perceived status and competition. Journal of Personality and Social Psychology, 82(6), 878–902. https://doi.org/10.1037/0022-3514.82.6.878
- Hu, L. T., & Bentler, P. M. (1999). Cutoff criteria for fit indexes in covariance structure analysis: Conventional criteria versus new alternatives. Structural Equation Modeling: A Multidisciplinary Journal, 6(1), 1–55. https://doi.org/10.1080/10705519909540118
- Li, X., Chan, K. W., & Kim, S. (2019). Service with emoticons: How customers interpret employee use of emoticons in online service encounters. Journal of Consumer Research, 45(5), 973–987. https://doi.org/10.1093/jcr/ucy016
- Spitzberg, B. H. (2000). What is good communication? Journal of the Association for Communication Administration, 29(1), 103–119.
13. Items of the Scale
Response Format:
7-point Likert scale (1 = Strongly disagree to 7 = Strongly agree) or 7-point semantic/differential scale
Scale Items:
- The employee’s behavior was appropriate.
- The employee’s behavior was acceptable.
- The employee’s behavior was proper.
- The employee’s behavior was fitting.
Scoring Protocol:
All items are positively keyed. Items are averaged to form an overall index of perceived appropriateness.