Abstract
The Classroom Assessment Fairness Inventory (CAFI), developed by Rasooli, DeLuca, Cheng, and Mousavi (2023), is an empirically validated psychometric instrument designed to capture and evaluate post-secondary and secondary students’ perceptions of fairness across multidimensional classroom assessment contexts. Drawing upon foundational theories of organizational justice, social psychology, and educational measurement, the CAFI operationalizes fairness not merely as technical accuracy or absence of statistical test bias, but as a socially constructed, situational judgment governed by principles of distributive, procedural, and interactional justice. The original inventory deploys a scenario-based methodology comprising five contextualized pedagogical vignettes—collaborative groupwork, summative examinations, academic integrity infractions (cheating), grading protocols, and formative feedback delivery—incorporating an initial pool of 40 discrete assessment actions aligned with theoretical fairness principles such as equity, equality, need, transparency, voice, bias suppression, correctability, justification, and respect.
Psychometric evaluation through exploratory factor analysis (EFA) and confirmatory factor analysis (CFA) with undergraduate cohorts confirmed a robust 21-item, five-factor structure: Unfairness in Groupwork, Fairness in Cheating, Fairness in Grading, Unfairness in Feedback, and Fairness in Feedback. The latent measurement model demonstrated acceptable to good goodness-of-fit indices: Root Mean Square Error of Approximation (RMSEA) = 0.06 (90% CI [0.05, 0.07]), Comparative Fit Index (CFI) = 0.91, and Tucker-Lewis Index (TLI) = 0.89. The subscales exhibited solid internal consistency reliability, with Cronbach’s alpha coefficients ranging from 0.79 to 0.87, complemented by a moderate inter-rater agreement coefficient of 0.56 (p < 0.001). Criterion-related and construct validity were substantiated via multivariate regression modeling against Dalbert’s Personal Belief in a Just World (PBJW) framework, demonstrating that dispositional justice beliefs significantly predict student fairness ratings across grading (β = 0.262, p < 0.05) and feedback (β = 0.294, p < 0.05). The CAFI represents an indispensable diagnostic and empirical tool for educational researchers, instructional designers, psychometricians, and classroom practitioners seeking to establish equitable, transparent, and defensible evaluative ecologies.
Keywords
Classroom Assessment Fairness, Student Perceptions, Unfairness in Groupwork, Fairness in Cheating, Fairness in Grading, Unfairness in Feedback, Fairness in Feedback, Educational Measurement, Organizational Justice, Classroom Climate, Procedural Justice, Distributive Justice
Authors
The Classroom Assessment Fairness Inventory was conceptualized, operationalized, and psychometrically validated by a collaborative team of measurement and assessment scholars in Canada:
- Amirhossein Rasooli (Lead Author and Investigator) — Killam Postdoctoral Fellow, Faculty of Education, University of Alberta, 11210 87 Ave NW, Edmonton, Alberta, Canada, T6G 2G5. Email: [email protected].
- Christopher DeLuca — Professor of Classroom Assessment and Associate Dean of Graduate Studies, Faculty of Education, Queen’s University, Kingston, Ontario, Canada.
- Liying Cheng — Professor and Director of the Assessment and Evaluation Group (AEG), Faculty of Education, Queen’s University, Kingston, Ontario, Canada. ORCID: 0000-0002-4458-5085.
- Amin Mousavi — Associate Professor of Educational Measurement and Psychometrics, Department of Educational Psychology and Special Education, College of Education, University of Saskatchewan, Saskatoon, Saskatchewan, Canada. ORCID: 0000-0002-6920-2319.
Purpose
The primary purpose of the Classroom Assessment Fairness Inventory (CAFI) is to provide an empirically grounded, multidimensional diagnostic instrument to systematically evaluate how students conceptualize, perceive, and weigh fairness across authentic classroom assessment practices. Historically, fairness within educational measurement has been conceptualized through a psychometric lens focused on standardized testing properties, such as Differential Item Functioning (DIF), construct-irrelevant variance, and predictive equivalence across demographic groups (e.g., the AERA, APA, & NCME Standards for Educational and Psychological Testing). However, within classroom-based, dynamic assessment ecologies, fairness is intrinsically interpersonal, experiential, and procedural. Traditional psychometric definitions fail to capture how everyday evaluative events—such as teacher grading adjustments, penalizing collaborative contributions, mitigating academic dishonesty, and delivering formative critiques—are mediated by students’ emotional, cognitive, and motivational appraisals.
The CAFI bridges this substantive gap by operationalizing assessment fairness through situational vignettes reflecting everyday pedagogical dilemmas. In high-stakes and low-stakes learning environments alike, perceived unfairness triggers negative educational sequelae, including cognitive disengagement, academic cynicism, diminished self-efficacy, communicative withdrawal, and elevated probabilities of retaliatory cheating or attrition. Conversely, when assessment practices are perceived as procedurally transparent, equitably distributed, and respectfully communicated, students exhibit elevated intrinsic motivation, heightened self-regulation, and deeper resilience in responding to critical academic feedback.
In research contexts, the CAFI serves as an empirical instrument to investigate how structural, instructional, and individual difference variables intersect to shape student perceptions of evaluative legitimacy. Researchers can deploy the CAFI to track institutional equity interventions, inspect demographic differences in fairness appraisals across race, gender, disability, and socioeconomic standing, and examine reciprocal associations with academic achievement, teacher-student alliance, and classroom socio-emotional climate. In instructional and institutional settings, faculty developers and higher education administrators can utilize the inventory as a professional development audit to identify blind spots in departmental grading policies, collaborative assessment designs, and feedback turnaround protocols.
Psychological Construct
The CAFI measures the psychological construct of perceived classroom assessment fairness, defined as a student’s cognitive appraisal and moral-affective evaluation of the justice, ethical integrity, and equity embedded within an educator’s assessment policies, evaluative interactions, and grading actions. Factor-analytic validation reduced the theoretical domain into a specialized five-factor measurement model that captures distinct yet complementary dimensions of justice appraisal across both positive (fairness) and negative (unfairness) valences:
1. Unfairness in Groupwork
This dimension taps into students’ acute sensitivity toward violations of distributive and procedural justice in collaborative peer learning environments. It measures appraisals of practices that enforce outcome equality without regard for heterogeneous individual contributions—colloquially recognized as the "free-rider problem." Key indicators evaluate educators assigning identical grades across all group members, denying students agency or choice in group composition, failing to detail evaluative rubrics before collaborative work begins, and dismissing student grievances regarding dysfunctional group dynamics. High scores reflect an acute perception of structural injustice when individual accountability is subsumed by collective grading mandates.
2. Fairness in Cheating
This subscale captures students’ endorsement of retributive justice, ethical integrity, and deterrence within the governance of academic dishonesty. Far from viewing strict sanctions as inherently hostile, students assess the fair application of moral consequences against individuals who compromise collective equity. It captures the appraisal of assigning zero marks for deliberate cheating, publicly reaffirming academic integrity boundaries, and demonstrating consistent disciplinary resolve, weighed against the necessity of allowing the accused student fair procedural voice and contextual justification.
3. Fairness in Grading
This construct assesses students’ endorsement of procedural consistency, bias suppression, need-based accommodations, and achievement-based distributive allocation. It measures perceptions of equity when teachers balance standardized achievement criteria (e.g., performance on transparently weighted summative examinations) with compassionate or individualized adjustments. Specifically, it captures how students view instructor interventions that consider extraneous hardships—such as adjusting final thresholds for failing, socioeconomically vulnerable, or at-risk students—versus practices that penalize academic marks for non-academic misconduct.
4. Unfairness in Feedback
Capturing the relational and interactional justice deficits in formative assessment cycles, this factor operationalizes the psychological harm, perceived favoritism, and neglect that occur during evaluative critique. It measures perceived violations when instructors direct detailed, actionable qualitative feedback exclusively toward high-performing or favored students while withholding formative commentary from struggling learners. It also captures the communicative indignity of disrespectful, dismissive, or arbitrary explanations regarding how feedback corresponds to student effort.
5. Fairness in Feedback
Complementary to the negative feedback valence, this factor isolates constructive interactional and procedural justice within formative assessment. It measures student appreciation for instructors who provide transparent, standardized evaluative rubrics prior to submission, guarantee rapid turnaround times (e.g., returning graded essays within four days), create explicit procedural avenues for post-assessment dialogue and clarification (voice), and calibrate feedback meaningfully against student investment and effort.
Theoretical Framework
The architectural foundation of the CAFI synthesizes classical organizational justice theory with modern socio-cognitive models of classroom assessment. For decades, justice scholars have categorized human fairness evaluations into three core pillars: distributive justice, procedural justice, and interactional justice. Rasooli et al. (2023) mapped these theoretical domains directly onto the interactive, pedagogical realities of classroom assessment ecologies.
Distributive justice, rooted in Adams’ Equity Theory (1965) and Deutsch’s social justice typologies (1975), addresses the allocation of outcomes—in this case, grades, marks, and academic credentials. Adams posited that individuals evaluate fairness by computing the ratio of their personal inputs (effort, time, cognitive engagement) to obtained outputs (grades, feedback), comparing this ratio against reference peers. In classroom assessment, distributive fairness operates across three competing allocation rules: equity (grades proportional to demonstrable individual contribution and competence), equality (every student receiving the exact same baseline mark, common in collaborative projects), and need (adjusting grades or offering remedial testing opportunities based on socio-emotional or systemic adversity).
Procedural justice, as formulated by Leventhal (1980) and further advanced by Thibaut and Walker (1975), posits that individuals evaluate institutions not merely by outcomes, but by the fairness of the processes used to determine those outcomes. Leventhal articulated six procedural justice criteria that Rasooli et al. embedded into the CAFI item specifications:
- Consistency: Assessment protocols must be maintained uniformly across all students and temporal occasions.
- Bias Suppression: The evaluator must remain free of personal favoritism, prejudice, or extraneous personal incentives.
- Accuracy: Evaluative judgments must reflect technically sound, valid measurement evidence.
- Correctability: Mechanisms, formal appeals, or grievance procedures must exist to rectify inaccurate grades.
- Representativeness (Voice): Students must be afforded voice, participatory choice, and consultative agency.
- Ethicality: Evaluative systems must align with prevailing professional, moral, and human standards.
Interactional justice, subdivided by Bies and Moag (1986) into interpersonal justice and informational justice, addresses the qualitative, human dimension of assessment delivery. Interpersonal justice dictates that instructors communicate evaluative decisions with dignity, empathy, and professional politeness, precluding sarcasm, public humiliation, or aggressive dismissal of student inquiries. Informational justice demands adequate, timely, and candid transparency regarding grading rubrics, syllabus weighting, assessment scheduling, and the underlying reasoning for specific evaluative decisions.
Finally, the CAFI framework intersects with Lerner’s (1980) Just World Theory, which asserts that individuals possess an implicit, fundamental need to believe that the world is an orderly, morally coherent system where individuals receive what they rightfully deserve. Students with strong personal beliefs in a just world (PBJW) possess an intrinsic cognitive schema that colors their interpretation of authority actions, predisposing them to perceive institutional grading and feedback rituals as legitimate, structured, and inherently fair, rather than arbitrary or hostile.
Validity
Validation of the CAFI was conducted through a rigorous multi-phase psychometric design conforming to the AERA/APA/NCME Standards. Initial content and face validity were established through systematic expert panels composed of assessment researchers, psychometricians, and classroom educators who reviewed the 40 original action statements across the five authentic scenarios. Items were evaluated for their contextual fidelity, construct representation, and theoretical alignment with justice dimensions (equity, voice, bias suppression, need, consistency, justification, and respect).
Construct validity was evaluated through structural equation modeling, incorporating both exploratory factor analysis (EFA) to extract latent dimensions and confirmatory factor analysis (CFA) on split validation samples of first-year undergraduate students in Canada (N spanning university cohorts aged 15 to 21+ years). The empirically derived five-factor model demonstrated exceptional structural integrity, successfully partitioning fairness attitudes into domain-specific appraisals rather than a monolithic global construct.
Criterion-related and convergent validity were established through multivariate regression models using Dalbert’s Personal Belief in a Just World (PBJW) scale as a theoretically grounded external criterion. In line with psychometric hypotheses:
- Personal Belief in a Just World significantly and positively predicted students’ perceptions of Fairness in Grading (β = 0.262, t = 3.41, p < 0.05).
- Personal Belief in a Just World significantly and positively predicted perceptions of Fairness in Feedback (β = 0.294, t = 3.88, p < 0.05).
- Conversely, PBJW emerged as a statistically significant negative predictor of perceived Unfairness in Feedback (β = −0.190, t = −2.45, p < 0.05).
These empirical trajectories confirm that students who maintain an internalized cognitive belief in a fair, predictable social order interpret normative teacher grading, timely feedback, and structured evaluative communications as manifestations of systemic equity, while exhibiting lower sensitivity to interpersonal feedback slights. Discriminant validity was supported by the distinct factor loadings and low-to-moderate inter-factor correlations among the five latent dimensions, confirming that students clearly distinguish between the fairness demands of collaborative peer work (groupwork), punitive academic policies (cheating), summative evaluation (grading), and formative communication (feedback).
Reliability
The reliability of the CAFI has been evaluated via internal consistency analysis and inter-rater agreement metrics, establishing that the instrument yields stable and psychometrically defensible scores across varied educational applications:
- Internal Consistency: Cronbach’s alpha (α) was computed for each of the five extracted latent factors. Across the factors, alpha coefficients consistently ranged from 0.79 to 0.87, satisfying conventional psychometric criteria for exploratory and applied assessment tools (where α ≥ 0.70 is deemed acceptable and α ≥ 0.80 is recognized as good). The subscales demonstrated robust homogeneity among their constitutive items, indicating that items within each vignette coherently tap into unified dimensions of justice appraisal without excessive redundancy.
- Inter-Rater Reliability: During the item coding and classification stages, where independent expert raters mapped the contextualized assessment actions to their theoretical underlying principles (e.g., Equity, Need, Voice, Bias Suppression), the overall inter-rater agreement demonstrated a statistically significant moderate average coefficient of 0.56 (p < 0.001). This empirical baseline confirms that while fairness scenarios evoke nuanced individual interpretive variance, the underlying justice constructs possess demonstrable objective cross-rater stability.
- Measurement Precision: Standard errors of measurement (SEM) across the subscales remain minimal, supporting the instrument’s capacity to discriminate between varying degrees of student justice orientations in large-scale educational research and programmatic evaluations.
Factor Analysis
The structural composition of the CAFI was analyzed through a two-stage factor-analytic framework utilizing both Exploratory Factor Analysis (EFA) and Confirmatory Factor Analysis (CFA). In the initial EFA phase, principal axis factoring with oblique rotation (promax) was conducted on the full 40-item pool to allow for natural intercorrelations among latent fairness constructs. Parallel analysis and scree plot inspections initially indicated potential models containing between five and six factors.
Subsequent Confirmatory Factor Analysis (CFA) was conducted using maximum likelihood estimation to rigorously test and compare the competing structural models:
| Tested Structural Model | RMSEA [90% CI] | CFI | TLI | Psychometric Disposition |
|---|---|---|---|---|
| Five-Factor Model (21 Retained Items) | 0.06 [0.05 – 0.07] | 0.91 | 0.89 | Retained as Best-Fitting & Conceptually Sound |
| Six-Factor Model | 0.07 [0.06 – 0.08] | 0.90 | 0.87 | Rejected due to Under-Identification (2 items) |
While the six-factor model yielded marginally acceptable statistical fit indices, psychometric inspection revealed that the sixth extracted factor retained only two items from Scenario 2 (Exam Actions S2-6 and S2-7: Mr. Ahmed not removing untaught questions, and Mr. Ahmed responding harshly to student complaints). In structural psychometrics, factors with fewer than three indicators are vulnerable to structural instability, under-identification, and conceptual inadequacy. Consequently, the authors eliminated the under-identified factor, consolidating the inventory into the definitive 21-item, five-factor model. All retained standardized item factor loadings in the final five-factor model were statistically significant (p < 0.001) and substantial, ranging between 0.48 and 0.84, confirming strong structural convergence.
Instrument / Measurement Tool
- Test Type: Original psychometric inventory utilizing situational, scenario-based evaluative items.
- Administration Format: Self-administered paper-and-pencil questionnaire or computerized online survey tool.
- Target Population: Secondary education students (ages 15–17) and post-secondary undergraduate and graduate students (ages 18–29 and older). Validated across diverse gender identities (male, female, non-binary).
- Total Number of Items: 21 final retained items situated within 5 overarching classroom pedagogical scenarios (with an initial exploratory pool of 40 actions across 5 domains).
- Structural Subscales:
- Subscale 1: Unfairness in Groupwork
- Subscale 2: Fairness in Cheating
- Subscale 3: Fairness in Grading
- Subscale 4: Unfairness in Feedback
- Subscale 5: Fairness in Feedback
- Response Format: A 6-point anchored Likert-type rating scale ranging from 1 ("Highly unfair") to 6 ("Highly fair"), accompanied by an explicit, separate non-scalar option: 7 = "I am not sure".
- Scoring and Transformation Procedures:
- Ratings for scalar items are scored linearly from 1 to 6.
- Selections of "7 = I am not sure" represent an explicit epistemic state of uncertainty; in empirical scoring, these responses are treated as missing data or dummy-coded separately for specific categorical cognitive response analyses, precluding distortion of metric means.
- Composite subscale scores are computed by calculating the arithmetic mean or sum of the items corresponding to each latent dimension. Separate scores for the fairness and unfairness valences are retained rather than summed into a single bipolar score, preserving unique perceptual variance.
- Estimated Completion Time: Approximately 15 to 25 minutes for reading the contextual vignettes and evaluating all associated action statements.
Permissions & Fee and Test Year
The Classroom Assessment Fairness Inventory (CAFI) was formally published in 2023 by Amirhossein Rasooli, Christopher DeLuca, Liying Cheng, and Amin Mousavi in the peer-reviewed journal Assessment in Education: Principles, Policy & Practice. The instrument is intended for non-commercial educational, diagnostic, and academic research purposes.
Fee Structure: There are no licensing or administration fees associated with using the CAFI for non-commercial academic research or scholarly investigations. However, formal copyright is held by the publisher (Taylor & Francis) and the contributing authors. Researchers, institutions, or practitioners planning to adapt, reproduce in full volumes, or digitize the scale for institutional deployment must contact the corresponding author, Dr. Amirhossein Rasooli ([email protected]), or request appropriate permissions through the journal publisher’s clearance center.
References
- Adams, J. S. (1965). Inequity in social exchange. In L. Berkowitz (Ed.), Advances in Experimental Social Psychology (Vol. 2, pp. 267–299). Academic Press. https://doi.org/10.1016/S0065-2601(08)60108-2
- Bies, R. J., & Moag, J. F. (1986). Interactional justice: Communication criteria of fairness. In R. J. Lewicki, B. H. Sheppard, & M. H. Bazerman (Eds.), Research on Negotiation in Organizations (Vol. 1, pp. 43–55). JAI Press.
- Dalbert, C. (1999). The world is more just for me than generally: Considering the personal belief in a just world scale’s validity and generality. Social Justice Research, 12(2), 79–98. https://doi.org/10.1023/A:1022091609047
- Deutsch, M. (1975). Equity, equality, and need: What determines which value will be used as the basis of distributive justice? Journal of Social Issues, 31(3), 137–149. https://doi.org/10.1111/j.1540-4560.1975.tb01000.x
- Lerner, M. J. (1980). The Belief in a Just World: A Fundamental Delusion. Plenum Press. https://doi.org/10.1007/978-1-4684-3602-0
- Leventhal, G. S. (1980). What should be done with equity theory? New approaches to the study of fairness in social relationships. In K. J. Gergen, M. S. Greenberg, & R. H. Willis (Eds.), Social Exchange: Advances in Theory and Research (pp. 27–55). Plenum Press. https://doi.org/10.1007/978-1-4613-3087-5_2
- Rasooli, A., DeLuca, C., Cheng, L., & Mousavi, A. (2023). Classroom assessment fairness inventory: A new instrument to support perceived fairness in classroom assessment. Assessment in Education: Principles, Policy & Practice, 30(5-6), 372–395. https://doi.org/10.1080/0969594X.2023.2255936
- Thibaut, J. W., & Walker, L. (1975). Procedural Justice: A Psychological Analysis. L. Erlbaum Associates.
- Tierney, R. D. (2014). Fairness as a multifaceted quality in classroom assessment. Studies in Educational Evaluation, 43, 55–69. https://doi.org/10.1016/j.stueduc.2013.12.003