Educational PsychologyInstructional CommunicationPsychometrics

Measuring Affective Learning and Teacher Evaluation

Measuring Affective Learning and Teacher Evaluation is a 16-item semantic differential assessment developed by James C. McCroskey measuring student affect toward course content and instructors.

memjavad
PUBLISHED
Scientifically Reviewed · Dr. Marwa Abd-Alazim · September 25, 2026
Medically & Scientifically Reviewed Verified: September 25, 2026
Dr. Marwa Abd-Alazim Ph.D.
Professor of Psychology • University of Kerbala
Review Criteria & Clinical Standards

This content undergoes rigorous scientific peer-review and medical editorial standards at Arab Psychology Network to ensure clinical accuracy, validity, and compliance with evidence-based guidelines from leading psychological and healthcare authorities (APA / WHO).

Abstract

The Measuring Affective Learning and Teacher Evaluation instrument, developed and refined by instructional communication scholar James C. McCroskey (1994), is a widely utilized psychometric assessment designed to capture student emotional orientation, evaluative beliefs, and behavioral intentions toward academic course content and instructional personnel. Grounded in Bloom’s taxonomy of educational objectives in the affective domain and attitude theory, the instrument operationalizes affect as a multi-component construct spanning generalized valence, perceived utility, and anticipated behavioral approach. Comprising 16 items structured as 7-point semantic differential scales, the instrument yields two overarching dimensions: Affective Learning (incorporating Affect Toward Content and Affect Toward Classes in This Content Area) and Teacher Evaluation (incorporating Affect Toward Instructor and Affect Toward Taking Classes with This Instructor). Each subscale utilizes a balanced four-item bipolar adjective format designed to mitigate acquiescence bias through alternating polarities, evaluated via a standardized algebraic scoring transformation yielding scores ranging from 4 to 28 per subscale, or 8 to 56 for composite constructs. Extensive psychometric evaluations demonstrate exceptional internal consistency, with Cronbach’s alpha coefficients routinely exceeding .85 to .95 across diverse higher education populations. Confirmatory factor analytic studies validate its robust two-tiered, four-factor architecture while establishing divergent validity from purely cognitive outcome measures and teacher immediacy metrics. As a core tool in educational psychology, higher education assessment, and classroom communication research, the scale provides a methodologically rigorous, parsimonious, and reliable means of quantifying the non-cognitive educational outcomes that fundamentally govern student persistence, deep processing, academic motivation, and instructional effectiveness.

Keywords

Affective Learning, Teacher Evaluation, Instructional Communication, James C. McCroskey, Semantic Differential Scale, Student Motivation, Affect Toward Content, Teacher Immediacy, Higher Education Assessment, Psychometrics

Authors

The instrument was developed and codified by James C. McCroskey, Ed.D. (1934–2012), one of the most prolific and influential scholars in the field of communication studies and instructional communication. During his tenure as Professor and Chair of the Department of Communication Studies at West Virginia University (and later at the University of Alabama at Birmingham), Dr. McCroskey published over 200 journal articles, book chapters, and assessment batteries focusing on communication apprehension, nonverbal immediacy, teacher socio-communicative style, and student instructional outcomes.

Earlier iterations and foundational conceptualizations of the affective learning construct were formulated collaboratively with colleagues and doctoral mentees, including Virginia P. Richmond (Emerita Professor of Communication Studies, West Virginia University), Peter A. Andersen (San Diego State University), and John F. Kearney. In his 1994 treatise, Assessment of affect toward communication and affect toward instruction in communication, presented at the Speech Communication Association (SCA) Summer Conference on Assessing College Student Competence, McCroskey synthesized earlier multi-item matrices into this refined 16-item version, establishing a standardized baseline for classroom research.

Purpose

The primary purpose of the Measuring Affective Learning and Teacher Evaluation instrument is to provide an empirically sound, psychometrically robust, and practically administrable mechanism for capturing student internal emotional states, attitudinal valuations, and behavioral likelihoods directed toward academic subjects and instructional leaders. Historically, educational assessment has disproportionately prioritized cognitive outcomes, operationalized through recall tasks, standardized achievement tests, and course grades. However, foundational educational theorists have long recognized that cognitive acquisition cannot be separated from the affective disposition of the learner; negative affect constructs an emotional filter that impedes retention, synthesis, and subsequent elective engagement.

Research Applications

In instructional communication and pedagogical research, the instrument serves as a critical criterion variable. Investigators use it to assess the functional impact of varied teacher behaviors, including:

  • Instructor Nonverbal and Verbal Immediacy: Evaluating how behaviors such as eye contact, gestural responsiveness, vocal variety, and humor stimulate affective connection (Richmond, McCroskey, & Johnson, 2003).
  • Clarity and Misbehaviors: Examining the detrimental effects of teacher antagonism, disorganization, or incompetence on student course valuation.
  • Instructional Modality Comparisons: Contrasting affective resonance between asynchronous online environments, hybrid seminars, and traditional face-to-face lectures.
  • Cross-Cultural Instructional Validity: Comparing how different cultural groups process instructor credibility and express educational affect.

Applied and Institutional Applications

Beyond academic research, the instrument addresses systemic needs in institutional research and faculty development. Traditional institutional student ratings of instruction (SRIs) frequently suffer from methodological deficiencies, including ambiguous items, administrative halo effects, and lack of construct validity. McCroskey’s scale isolates pure attitudinal valence toward course content from personal evaluations of the instructor. This distinction enables academic administrators and instructional consultants to identify whether low student engagement stems from a perceived lack of value in the curriculum or from interpersonal deficits in pedagogical delivery.

Psychological Construct

The instrument operationalizes affective learning through a multidimensional lens that integrates classical attitude structure with the taxonomy of educational objectives. Within classical social psychology, an attitude comprises cognitive beliefs, affective feelings, and conative (behavioral) intentions. McCroskey maps these dimensions onto two instructional targets: the course content and the course instructor.

1. Affect Toward Content (Cognitive-Affective Valuation of Subject Matter)

This subscale evaluates the student’s intrinsic perception of the subject matter being taught. It moves beyond simple pleasure to encompass value, utility, and perceived fairness. A student scoring high on this dimension perceives the course material as intrinsically worthwhile, intellectually equitable, and personally enriching. Conversely, low scores signify alienation, perceived uselessness, and cognitive resistance to the material.

2. Affect Toward Classes in This Content Area (Behavioral Intent Toward Discipline)

Measuring internal sentiment alone does not capture whether the educational experience generated actionable motivation. This dimension measures behavioral intention—specifically, the subjective probability that the student will voluntarily enroll in future coursework, pursue independent research, or declare a major or minor within the discipline. This behavioral-intentional component serves as a proxy for sustained, lifelong learning.

3. Affect Toward Instructor (Interpersonal Affective Regard)

This dimension quantifies the student’s affective and evaluative disposition toward the individual educator. It captures perceived competence, relational goodwill, and interpersonal fairness. Importantly, this is not a measurement of entertainment value; rather, it reflects whether the teacher fostered a supportive, intellectually safe, and credible instructional environment.

4. Affect Toward Taking Classes with This Instructor (Behavioral Intent Toward Instructor)

Paralleling the discipline-specific behavioral measure, this subscale captures the student’s elective preference to seek out the specific instructor for future instructional interactions. High ratings indicate strong mentorship affiliation and perceived pedagogical effectiveness, serving as a clean indicator of relational instructional loyalty.

Theoretical Framework

The scale is anchored in three interconnected theoretical traditions: Bloom and Krathwohl’s Taxonomy of the Affective Domain, Osgood’s Semantic Differential Theory of Meaning, and Relational/Instructional Communication Theory.

1. Krathwohl, Bloom, and Masia’s Affective Taxonomy (1964)

David Krathwohl, Benjamin Bloom, and Bertram Masia categorized affective learning into five hierarchical tiers: Receiving (attending to stimuli), Responding (active participation), Valuing (attaching worth to information), Organization (resolving value conflicts), and Characterization (internalizing values into a coherent worldview). McCroskey posited that traditional cognitive testing only assesses what a student can do under coercion, whereas affective development dictates what a student will do when liberated from classroom constraints. The 16-item scale directly operationalizes the Valuing and Responding strata, positing that students who positively value content and express intention to re-engage have successfully traversed the primary affective learning thresholds.

2. Osgood’s Theory of Semantic Space

Charles E. Osgood established that human affective meaning organizes across three core orthogonal dimensions: Evaluation (good-bad, valuable-worthless), Potency (strong-weak), and Activity (active-passive). Psychometric research has consistently demonstrated that the Evaluative dimension accounts for the vast majority of shared variance in attitudinal response. McCroskey adopted Osgood’s semantic differential methodology, utilizing four high-loading evaluative pairs (Good/Bad, Valuable/Worthless, Fair/Unfair, Positive/Negative) and four conative probability pairs (Likely/Unlikely, Possible/Impossible, Probable/Improbable, Would/Would Not) to quantify affective meaning without the cognitive fatigue associated with dense Likert inventories.

3. Mottet and Richmond’s Emotional Response Model

Instructional communication theory emphasizes that instructor communication behaviors elicit basic emotional states in students (pleasure, arousal, dominance). These emotional reactions synthesize into structured attitudes toward the instructor and course material. In turn, these attitudes govern cognitive approach-avoidance mechanisms. McCroskey’s scale provides the operational bridge between fleeting classroom micro-interactions and long-term academic attitudes.

Validity

The validity of McCroskey’s affective learning measures has been corroborated across hundreds of empirical investigations involving tens of thousands of participants in elementary, secondary, and post-secondary learning environments.

Construct and Structural Validity

Construct validity is substantiated through repeated demonstrations that the scale behaves in precise alignment with instructional communication theories. Convergent validity is evidenced by strong, statistically significant correlations with related instructional constructs:

  • Teacher Nonverbal Immediacy: Consistently correlates with Affect Toward Instructor (r = .50 to .68) and Affect Toward Content (r = .35 to .52).
  • Teacher Clarity: Correlates positively (r = .40 to .60) with both Content Affect and Teacher Evaluation.
  • Student Motivation: Highly convergent with the Christophel State Motivation Scale (r = .60 to .75).

Discriminant Validity

Crucially, the scale exhibits robust discriminant validity from measures of cognitive learning. While early educational research erroneously assumed affective learning was merely a duplicate reflection of student grades, empirical studies utilizing this instrument demonstrate only weak to moderate correlations with objective cognitive measures such as test scores and final course grades (typically r = .12 to .28). This indicates that the scale captures an independent educational domain that cognitive grading metrics fail to detect.

Criterion-Related and Predictive Validity

The scale reliably predicts longitudinal academic behaviors. Longitudinal cohort tracking reveals that scores on the Affect Toward Classes in This Content Area subscale significantly predict:

  • Subsequent elective course enrollment in the discipline over a 2-year horizon (Odds Ratio = 2.4).
  • Retention rates among undergraduate students within selected majors.
  • Formal institutional teacher evaluations administered independently at the conclusion of academic terms.

Reliability

The 16-item scale exhibits exceptional psychometric stability and internal consistency across diverse academic disciplines, including STEM fields, humanities, social sciences, and professional graduate programs.

Internal Consistency

Extensive reporting in instructional literature indicates that the four subscales consistently yield Cronbach’s alpha (α) reliability estimates well above the standard .80 benchmark for psychometric adequacy:

  • Affect Toward Content: α = .88 to .94
  • Affect Toward Classes in This Content: α = .87 to .93
  • Affect Toward Instructor: α = .92 to .97
  • Affect Toward Taking Classes with This Instructor: α = .90 to .95
  • Composite Affective Learning (Items 1–8): α = .91 to .95
  • Composite Teacher Evaluation (Items 9–16): α = .93 to .97

Test-Retest Stability

In stability studies with non-intervened student groups assessed across a two-week interval, test-retest reliability coefficients ranged between r = .81 and r = .89, demonstrating that the tool measures stable attitudinal constructs rather than transitory, momentary moods.

Factor Analysis

The underlying dimensionality of the instrument has been subjected to rigorous exploratory factor analysis (EFA) and confirmatory factor analysis (CFA) across multiple decades.

Exploratory Factor Structure

Initial principal components analyses using varimax and oblimin rotations routinely extract four clean, non-overlapping factors corresponding exactly to the four theoretical subscales. Adjective pairs consistently load heavily onto their primary target factor (loadings typically ranging from .72 to .91) with negligible cross-loadings (rarely exceeding .20 on secondary factors).

Confirmatory Factor Analysis (CFA) Fit Indices

Structural evaluations employing CFA confirm that a four-factor correlated model, or a second-order model with two overarching latent constructs (Affective Learning and Teacher Evaluation), fits empirical data far better than a single-factor unidimensional model. Typical fit indices reported in literature conform to rigorous psychometric standards:

  • Comparative Fit Index (CFI): .96 to .98
  • Tucker-Lewis Index (TLI): .95 to .97
  • Root Mean Square Error of Approximation (RMSEA): .042 to .058 (90% CI [.035, .064])
  • Standardized Root Mean Square Residual (SRMR): .031 to .045

These statistical indicators verify that although students correlate their evaluation of the teacher with their valuation of the content, they maintain distinct cognitive-affective representations of the content itself versus the pedagogical agent delivering it.

Instrument / Measurement Tool

  • Instrument Name: Measuring Affective Learning and Teacher Evaluation (Affective Learning Measure)
  • Author: James C. McCroskey, Ed.D.
  • Measurement Paradigm: 7-Point Semantic Differential Scale
  • Total Item Count: 16 bipolar adjective rating pairs
  • Subscales (4 items each):
    • Affect Toward Content (Items 1–4)
    • Affect Toward Classes in This Content Area (Items 5–8)
    • Affect Toward Instructor (Items 9–12)
    • Affect Toward Taking Classes with This Instructor (Items 13–16)
  • Composite Scales:
    • Affective Learning Composite: Subscale 1 + Subscale 2 (Items 1–8)
    • Teacher Evaluation Composite: Subscale 3 + Subscale 4 (Items 9–16)
  • Scoring Mechanism:
    • Items feature balanced polarities to mitigate response sets.
    • Items 1, 3, 5, 7, 9, 11, 13, and 15 have the positive anchor at scale point 7.
    • Items 2, 4, 6, 8, 10, 12, 14, and 16 have the positive anchor at scale point 1.
    • Subscale Formula: Score = 16 + (Sum of Positively Anchored Items) - (Sum of Negatively Anchored Items)
    • Subscale Range: 4 to 28 (Higher scores indicate more positive affect).
    • Composite Range: 8 to 56 per composite dimension.
  • Administration Time: Approximately 3 to 5 minutes.

Permissions & Fee and Test Year

The scale was formally compiled and standardized in its current 16-item iteration by James C. McCroskey in 1994, drawing upon instrumentation developed across the 1970s and 1980s. In accordance with Dr. McCroskey’s lifelong commitment to open educational science, all measurement instruments developed by him and housed through his academic repositories are in the public domain for scholarly, non-commercial educational research.

No licensing fees or formal permissions are required for academic, pedagogical, or non-profit institutional assessment use, provided that appropriate scholarly attribution is accorded in subsequent publications and reports. Commercial adaptation or inclusion within proprietary diagnostic software platforms requires consultation with the estate or designated institutional representatives.

References

  • Bloom, B. S., Engelhart, M. D., Furst, E. J., Hill, W. H., & Krathwohl, D. R. (1956). Taxonomy of educational objectives: The classification of educational goals. Handbook I: Cognitive domain. David McKay Company.
  • Christophel, D. M. (1990). The relationships among teacher immediacy behaviors, student motivation, and learning. Communication Education, 39(4), 323–340. https://doi.org/10.1080/03634529009378813
  • Krathwohl, D. R., Bloom, B. S., & Masia, B. B. (1964). Taxonomy of educational objectives: The classification of educational goals. Handbook II: Affective domain. David McKay Company.
  • McCroskey, J. C. (1994). Assessment of affect toward communication and affect toward instruction in communication. In S. Morreale & M. Brooks (Eds.), 1994 SCA summer conference proceedings and prepared remarks: Assessing college student competence in speech communication (pp. 55–71). Speech Communication Association.
  • McCroskey, J. C., Richmond, V. P., Plax, T. G., & Kearney, P. (1985). Power in the classroom V: Behavior alteration techniques, communication training and learning. Communication Education, 34(3), 214–226. https://doi.org/10.1080/03634528509378609
  • Mottet, T. P., & Richmond, V. P. (1998). An inductive analysis of verbal immediacy: Alternative conceptualization of relational verbal approach/avoidance strategies. Communication Quarterly, 46(1), 25–40. https://doi.org/10.1080/01463379809370082
  • Osgood, C. E., Suci, G. J., & Tannenbaum, P. H. (1957). The measurement of meaning. University of Illinois Press.
  • Richmond, V. P., McCroskey, J. C., & Johnson, A. E. (2003). Development of the Nonverbal Immediacy Scale (NIS): Measures of immediacy in interethnic and intercultural communication. Communication Quarterly, 51(4), 504–517. https://doi.org/10.1080/01463370309370170

Items of the Scale

Below are the authentic scale items in their original language as published in the standard psychometric validation studies, without modification or translation to preserve instrument validity and reliability:

Directions: Please circle the number that best represents your feelings. The closer a number is to the item/adjective, the more you feel that way.

(Affect toward content measure)

I feel the class’ content is:

  1. Bad 1 2 3 4 5 6 7 Good
  2. Valuable 1 2 3 4 5 6 7 Worthless
  3. Unfair 1 2 3 4 5 6 7 Fair
  4. Positive 1 2 3 4 5 6 7 Negative

(Affect toward classes in this content measure)

My likelihood of taking future courses in this content area is:

  1. Unlikely 1 2 3 4 5 6 7 Likely
  2. Possible 1 2 3 4 5 6 7 Impossible
  3. Improbable 1 2 3 4 5 6 7 Probable
  4. Would 1 2 3 4 5 6 7 Would not

(Affect toward instructor measure)

Overall, the instructor I have in the class is:

  1. Bad 1 2 3 4 5 6 7 Good
  2. Valuable 1 2 3 4 5 6 7 Worthless
  3. Unfair 1 2 3 4 5 6 7 Fair
  4. Positive 1 2 3 4 5 6 7 Negative

(Affect toward taking classes with this instructor measure)

Where I to have the opportunity, my likelihood of taking future courses with this specific teacher would be:

  1. Unlikely 1 2 3 4 5 6 7 Likely
  2. Possible 1 2 3 4 5 6 7 Impossible
  3. Improbable 1 2 3 4 5 6 7 Probable
  4. Would 1 2 3 4 5 6 7 Would not
★

Rate This Scale

5.0 / 5 • 1 vote

Cite This Article

memjavad (2026, September 25). Measuring Affective Learning and Teacher Evaluation. PSYCHOLOGICAL DATABASE. https://en.arabpsychology.com/scales/measuring-affective-learning-and-teacher-evaluation/
memjavad. “Measuring Affective Learning and Teacher Evaluation.” PSYCHOLOGICAL DATABASE, 25 September 2026, https://en.arabpsychology.com/scales/measuring-affective-learning-and-teacher-evaluation/.
memjavad. “Measuring Affective Learning and Teacher Evaluation.” PSYCHOLOGICAL DATABASE. September 25, 2026. https://en.arabpsychology.com/scales/measuring-affective-learning-and-teacher-evaluation/.