1. Abstract
The Student Satisfaction with Teaching Scale (SSTS) is a multidimensional psychometric instrument designed to assess tertiary students’ evaluations and perceived satisfaction with higher education instructional quality. Derived from the conceptual and empirical architecture of Herbert W. Marsh’s Students’ Evaluations of Educational Quality (SEEQ) framework, the SSTS operationalizes instructional effectiveness across five primary pedagogical dimensions: Perceived Learning Value, Instructor Enthusiasm, Course Organisation, Group Interaction Quality, and Individual Rapport. The scale typically comprises 20 to 25 items measured on a 5-point or 9-point Likert-type response format, ranging from strongly disagree (or very poor) to strongly agree (or very good). Extensive psychometric investigations indicate that the SSTS exhibits exceptional construct validity, supported by confirmatory factor analyses demonstrating robust goodness-of-fit indices (CFI > .95, TLI > .94, RMSEA < .05). Internal consistency reliability across the five subscales is consistently high, with Cronbach’s alpha coefficients routinely exceeding .85 and McDonald’s omega hierarchical estimates surpassing .80. The scale displays strong convergent validity with student cognitive engagement and objective course performance metrics, alongside discriminant validity against extraneous variables such as class size and elective versus mandatory course status. In contemporary tertiary settings, the SSTS serves as a vital diagnostic and empirical tool for faculty formative self-assessment, institutional quality assurance, and the broader Scholarship of Teaching and Learning (SoTL).
2. Keywords
Student Satisfaction with Teaching Scale, SEEQ, student evaluations of teaching, instructional effectiveness, higher education pedagogy, psychometrics, learning value, instructor enthusiasm, course organization, group interaction, teacher-student rapport, confirmatory factor analysis.
3. Authors
The conceptual origin of the instrument rests upon the foundational psychometric work of Herbert W. Marsh (Distinguished Professor of Educational Psychology, Self-concept Enhancement and Learning Facilitation [SELF] Research Centre, Western Sydney University, Australia; formerly Department of Education, University of Oxford, United Kingdom). Contemporary adaptations, structural validations, and psychometric refinements of the five-factor Student Satisfaction with Teaching Scale have been advanced by diverse psychometricians and educational evaluation researchers globally within university assessment and quality-enhancement consortia.
4. Purpose
The primary purpose of the Student Satisfaction with Teaching Scale is to provide an empirically defensible, multidimensional measurement of instructional quality as perceived by university students. For decades, tertiary institutions relied on unidimensional global rating items (e.g., "Overall, this instructor was effective"), which conflated distinct pedagogical competencies, yielded susceptibility to halo effects, and offered virtually no actionable diagnostic feedback for pedagogical improvement.
The SSTS addresses these limitations by disaggregating instructional quality into distinct, actionable pedagogical domains. The theoretical and empirical rationale for this approach is grounded in the recognition that university teaching is inherently multifaceted; an instructor may demonstrate exceptional organizational precision and conceptual clarity while exhibiting limited spontaneous enthusiasm or conversational rapport. By capturing these nuances, the SSTS fulfills several institutional and scholarly objectives:
- Formative Instructional Improvement: Providing academic staff with granular, domain-specific diagnostic data that highlight pedagogical strengths and concrete targets for instructional redesign.
- Summative Evaluation and Institutional Benchmarking: Supplying department chairs, tenure committees, and accreditation bodies with psychometrically sound metrics that withstand administrative scrutiny and minimize systematic evaluation bias.
- Scholarship of Teaching and Learning (SoTL): Serving as a rigorous dependent variable in quasi-experimental and longitudinal educational research evaluating the efficacy of pedagogical interventions, active learning methodologies, and digital learning transformations.
- Student Agency and Feedback Loops: Enhancing learner engagement and academic self-efficacy by providing transparent, structured mechanisms through which student voices directly inform the instructional design of the academic curriculum.
5. Psychological Construct
The Student Satisfaction with Teaching Scale measures a multidimensional psychological construct rooted in educational psychology and social-cognitive communication models. Instructional effectiveness and perceived satisfaction are conceptualized not as an innate teacher personality trait, but as an interactive transactional phenomenon occurring between instructor behaviors, course architecture, and student cognitive processing. The five core dimensions measured by the SSTS include:
1. Perceived Learning Value
This subscale evaluates students’ subjective cognitive appraisal of their academic growth, intellectual stimulation, and skill acquisition during the course. Rather than measuring rote memorization, it captures the extent to which the course broadened student perspectives, stimulated critical inquiry, and generated intrinsic intellectual value. Exemplary item content targets feelings of personal accomplishment, deeper mastery of complex subject matter, and perceived real-world relevance.
2. Instructor Enthusiasm
Instructor Enthusiasm operationalizes the dynamic nonverbal and verbal expressive behaviors of the educator, including vocal inflection, animated physical delivery, manifest passion for the subject matter, and the capacity to generate sustained student curiosity. Grounded in research on cognitive contagion and teacher immediacy, this construct assesses whether the instructor’s personal engagement fosters a stimulating learning climate that mitigates academic boredom and cognitive fatigue.
3. Course Organisation
Course Organisation measures the structural predictability, coherence, and instructional alignment of the curricular unit. This dimension includes the clarity of course objectives, systematic sequencing of modular content, explicit alignment between syllabus expectations and evaluation criteria, and the timely provision of educational resources. High scores on this subscale reflect pedagogical scaffolding that minimizes extraneous cognitive load and allows students to allocate mental resources directly to core conceptual learning.
4. Group Interaction Quality
Rooted in social constructivist educational theory, Group Interaction Quality examines the participatory social architecture of the lecture hall, seminar, or laboratory. It assesses the instructor’s ability to facilitate student discussions, encourage differing viewpoints, cultivate an emotionally safe space for intellectual risk-taking, and foster reciprocal academic dialogue among peers. This dimension captures whether the learning environment is dialogic or exclusively didactic.
5. Individual Rapport
Individual Rapport reflects the perceived interpersonal accessibility, empathy, and professional approachability of the instructor. It assesses the instructor’s genuine concern for individual student development, availability outside scheduled contact hours, non-judgmental guidance during academic difficulty, and fairness in student interactions. This construct operationalizes the relational dynamic that underpins student academic resilience and belonging within higher education settings.
6. Theoretical Framework
The theoretical foundations of the SSTS integrate Social Constructivism, Cognitive Load Theory, and Self-Determination Theory (SDT). The foundational psychometric architecture directly builds upon Herbert W. Marsh’s multidimensional model of student evaluations of educational quality. Historically, educational theorists debate whether instructional quality can be captured through a single general factor (a "g-factor" of teaching effectiveness) or requires a distinct multidimensional structure. Marsh (1982, 1984) empirically established that student evaluations are distinctly multidimensional, demonstrating that distinct components of teaching are differentially correlated with external criteria such as objective student learning, subsequent course enrollment, and instructor self-evaluations.
From the perspective of Self-Determination Theory (Deci & Ryan, 2000), the five dimensions directly map onto the satisfaction of basic psychological needs:
- Autonomy and Competence: Facilitated through Perceived Learning Value and clear Course Organisation, providing students with structured roadmaps to mastery and cognitive self-regulation.
- Relatedness: Supported through Individual Rapport and Group Interaction Quality, which foster a supportive community of inquiry where students feel recognized, respected, and socially integrated.
- Intrinsic Motivation: Sparked and sustained by Instructor Enthusiasm, which functions as an emotional scaffold transforming dry conceptual paradigms into engaging intellectual challenges.
Furthermore, Cognitive Load Theory (Sweller, 2011) underpins the Course Organisation dimension: when course materials, schedules, and learning outcomes are poorly structured, students experience high extraneous cognitive load, impeding germane schema acquisition. By systematically assessing these structural elements, the SSTS bridges cognitive architecture and relational pedagogy.
7. Validity
The psychometric validity of the SSTS and its parent SEEQ framework has been established across thousands of higher education courses spanning diverse disciplines, including STEM, humanities, social sciences, and professional degree programs.
Construct Validity
Construct validity has been extensively confirmed through both within-network and between-network approaches. Large-scale structural equation modeling consistently supports the five correlated first-order factors over unidimensional alternatives. Studies examining goodness-of-fit metrics across multi-institution samples typically document Tucker-Lewis Index (TLI) and Comparative Fit Index (CFI) values ranging between .94 and .98, with Root Mean Square Error of Approximation (RMSEA) values settling comfortably below .05 (Marsh, 1987; Marsh & Roche, 1997).
Convergent Validity
The scale demonstrates substantial convergent validity when correlated with external educational benchmarks:
- Multi-Section Validity Studies: In multisection courses with standardized identical final examinations, sections that assign higher ratings to Perceived Learning Value and Course Organisation demonstrate statistically significant higher average final exam scores (correlations ranging from r = .30 to r = .52, p < .001).
- Instructor Self-Evaluations: Studies administering parallel rating inventories to instructors themselves indicate substantial agreement between student ratings on the SSTS and faculty self-ratings, with convergent correlations across matching dimensions averaging r = .45 to .65, while divergent correlations remain significantly lower (r < .20).
- Peer and Administrator Observations: Independent ratings conducted by trained peer observers who observe classroom sessions correlate moderately to strongly with student aggregated scores on Instructor Enthusiasm and Course Organisation (r = .40–.58).
Discriminant Validity
A critical concern in student evaluation methodology involves potential contamination by extraneous confounding variables. Psychometric evaluations demonstrate that SSTS subscale scores are largely unaffected by factors such as class size, student expected grades (when prior grade point average is controlled), time of day, or instructor gender, displaying minimal variance attribution (typically < 5% of total variance explained by non-instructional background variables).
8. Reliability
The SSTS demonstrates exceptional empirical reliability across both individual student responses and class-aggregated evaluation metrics. Educational measurement theory distinguishes between the student as the unit of analysis and the classroom section as the unit of analysis; the SSTS excels at both levels.
Internal Consistency
Internal consistency metrics for the five subscales routinely surpass rigorous psychometric standards for educational decision-making. Typical empirical findings across literature reviews yield the following reliability parameters:
- Perceived Learning Value: Cronbach’s α = .88 – .93; McDonald’s ω = .89
- Instructor Enthusiasm: Cronbach’s α = .89 – .94; McDonald’s ω = .91
- Course Organisation: Cronbach’s α = .86 – .91; McDonald’s ω = .87
- Group Interaction Quality: Cronbach’s α = .84 – .90; McDonald’s ω = .86
- Individual Rapport: Cronbach’s α = .85 – .92; McDonald’s ω = .88
Test-Retest Stability and Generalizability
The scale demonstrates robust test-retest reliability across temporal intervals. When ratings are gathered midway through an academic semester and repeated at the conclusion of the term, stability coefficients range between r = .70 and .84. Longitudinal studies examining the same instructor teaching the same course across consecutive academic years demonstrate instructor stability correlations exceeding r = .80, indicating that the scale reliably captures stable instructional competencies rather than transient cohort effects.
Generalizability Theory (G-Theory) studies further show that with class sample sizes of 15 or more responding students, the generalizability coefficient (Φ-coefficient) routinely exceeds .90, confirming that class-aggregated scores provide an exceptionally dependable index of overall instructional quality.
9. Factor Analysis
The factorial validity of the SSTS has been rigorously scrutinized using both Exploratory Factor Analysis (EFA) and Confirmatory Factor Analysis (CFA) methodologies across decades of psychometric inquiry.
Exploratory Factor Analysis (EFA)
Early structural investigations utilizing principal axis factoring with oblique rotations (e.g., Promax, Direct Oblimin) invariably yield a clean, interpretable five-factor solution corresponding directly to the theoretical constructs. Eigenvalues for the five components reliably exceed the Kaiser-Guttman criterion (> 1.0), cumulatively accounting for approximately 65% to 75% of the total item variance. Pattern matrix loadings consistently reveal high primary factor loadings (λ = .65 to .88) with minimal cross-loadings (rarely exceeding .20).
Confirmatory Factor Analysis (CFA) and Model Fit
Confirmatory factor analyses testing competing structural architectures have demonstrated the definitive superiority of the oblique five-factor model over both single-factor (unidimensional) and orthogonal orthogonal models. Representative fit indices derived from robust maximum likelihood (MLR) estimation across multi-thousand student datasets exhibit exceptional precision:
- Comparative Fit Index (CFI): .962 to .978
- Tucker-Lewis Index (TLI): .955 to .971
- Root Mean Square Error of Approximation (RMSEA): .038 to .046 (90% CI [.034, .049])
- Standardized Root Mean Square Residual (SRMR): .031 to .039
- χ²/df Ratio: Consistently between 1.8 and 2.6
Higher-order CFA models specifying a general second-order factor ("Overall Teaching Effectiveness") accounting for the correlations among the five first-order dimensions have also demonstrated acceptable fit (CFI > .94, RMSEA < .055). However, hierarchical model comparisons via likelihood ratio tests consistently favor the multidimensional first-order model, demonstrating that substantial unique variance resides within each discrete pedagogical domain that would be obscured by reliance on a single composite score.
10. Instrument / Measurement Tool
- Instrument Name: Student Satisfaction with Teaching Scale (SSTS)
- Underlying Framework: Adapted from the Students’ Evaluations of Educational Quality (SEEQ) model (Marsh, 1982)
- Target Population: Undergraduate and postgraduate university students across all academic disciplines
- Administration Format: Standardized self-report inventory administered paper-and-pencil or electronically via learning management systems (LMS)
- Administration Time: Approximately 8 to 12 minutes
- Total Item Count: 20 to 25 items (typically 4 to 5 balanced items per dimension)
- Subscale Architecture:
- Perceived Learning Value (4–5 items)
- Instructor Enthusiasm (4–5 items)
- Course Organisation (4–5 items)
- Group Interaction Quality (4–5 items)
- Individual Rapport (4–5 items)
- Response Format: 5-point Likert scale (1 = Strongly Disagree, 2 = Disagree, 3 = Neutral / Undecided, 4 = Agree, 5 = Strongly Agree) or 9-point anchored evaluation scale (1 = Very Poor to 9 = Very Good)
- Scoring Protocol: Dimensional scores are computed by calculating the arithmetic mean of all completed items within each respective subscale. Scores range from 1.0 to 5.0 (or 1.0 to 9.0). Higher scores denote greater perceived instructional quality and higher student satisfaction. Negative or reverse-worded items (if included) are reverse-scored prior to aggregation.
11. Permissions & Fee and Test Year
The intellectual roots of the scale date back to Herbert W. Marsh’s seminal 1982 publication on the SEEQ instrument in the British Journal of Educational Psychology. Marsh’s original SEEQ items and their direct five-factor adaptations for academic research and institutional quality evaluation have historically been made widely accessible to the global academic community for non-commercial educational research, institutional review, and faculty development without royalty fees. However, commercial reproduction, integration into proprietary software suites, or corporate commercial distribution typically requires formal institutional permission or licensing from the copyright holders and publishers. Researchers and institutions utilizing the SSTS are strongly advised to cite the primary source literature appropriately and review any specific institutional copyright guidelines governing local adaptations.
12. References
Marsh, H. W. (1982). SEEQ: A reliable, valid, and useful instrument for collecting students’ evaluations of university teaching. British Journal of Educational Psychology, 52(1), 77–95. https://doi.org/10.1111/j.2044-8279.1982.tb02505.x
Marsh, H. W. (1984). Students’ evaluations of university teaching: Dimensionality, reliability, validity, potential baises, and utility. Journal of Educational Psychology, 76(5), 707–754. https://doi.org/10.1037/0022-0663.76.5.707
Marsh, H. W. (1987). Students’ evaluations of university teaching: Research findings, methodological issues, and directions for future research. International Journal of Educational Research, 11(3), 253–388. https://doi.org/10.1016/0883-0355(87)90001-2
Marsh, H. W., & Roche, L. A. (1997). Making students’ evaluations of teaching effectiveness effective: The critical issues of validity, bias, and utility. American Psychologist, 52(11), 1187–1197. https://doi.org/10.1037/0003-066X.52.11.1187
Deci, E. L., & Ryan, R. M. (2000). The "what" and "why" of goal pursuits: Human needs and the self-determination of behavior. Psychological Inquiry, 11(4), 227–268. https://doi.org/10.1207/S15327965PLI1104_01
Sweller, J. (2011). Cognitive load theory. In J. P. Mestre & B. H. Ross (Eds.), The psychology of learning and motivation: Cognition in education (Vol. 55, pp. 37–76). Academic Press. https://doi.org/10.1016/B978-0-12-387690-4.00002-8