Abstract
The Teacher Evaluation Research Survey (TERS) is an administrative and psychometric survey instrument developed by Mark E. Himmelein (2009) to examine public school principals’ supervisory practices, operational contexts, and professional attitudes toward teacher evaluation systems. Designed in the context of emerging statewide accountability reforms, the TERS captures the tension between summative accountability (administrative decision-making, contract renewal, remediation, dismissal) and formative supervision (pedagogical growth, instructional coaching, professional development). The instrument comprises a multi-part, mixed-format battery featuring demographic profiles, institutional operational characteristics, descriptive inventories of current district evaluation frameworks, evaluative opinions regarding procedural adequacy and revision needs, and multi-item Likert scales assessing principals’ support for diverse evaluation data sources (e.g., formal/informal observations, student achievement and value-added metrics, teacher and student portfolios, artifacts, and multi-rater stakeholder surveys). Furthermore, the TERS operationalizes comparative perceptual discrepancy by assessing principals’ attitudes toward evaluation components in formal vs. informal systems and contrasting those ratings against principals’ perceptions of teacher receptivity. The instrument’s psychometric foundation draws upon established personnel evaluation inventories, including the Teacher Evaluation Practices Survey (Loup et al., 1996) and the Teacher Evaluation Needs Identification Survey (Iwanicki, 1982), refined through expert review panels and superintendent focus groups. The attitudinal rating matrices employ a 5-point Likert response scale ranging from 1 (Strongly Against) to 5 (Strongly Favor). The TERS provides educational researchers, psychometricians, and school district administrators with an empirically rigorous diagnostic tool for evaluating administrative workload, evaluator readiness, collective bargaining alignment, and leadership attitudes toward modern instructional accountability systems.
Keywords
Teacher evaluation, principal attitudes, educational leadership, formative assessment, summative evaluation, psychometrics, instructional supervision, Charlotte Danielson framework, personnel appraisal, educational administration, value-added modeling, survey research.
Authors
The Teacher Evaluation Research Survey was conceived, developed, and validated by Dr. Mark E. Himmelein in 2009 as part of doctoral research in the Department of Educational Leadership within the Judith Herb College of Education at the University of Toledo, Ohio, United States. Dr. Himmelein’s work centered on educational administration, educational policy analysis, and the systemic organizational mechanisms governing teacher performance appraisal in Midwestern public school districts.
Purpose
The primary purpose of the Teacher Evaluation Research Survey (TERS) is to systematically capture, quantify, and analyze the operational structures of school-level teacher evaluation systems alongside the personal, administrative, and professional attitudes of practicing building principals toward those systems. Historically, teacher evaluation in American public education suffered from procedural superficiality—often characterized in organizational literature as “the widget effect” (Weisberg et al., 2009)—wherein virtually all teachers were rated as satisfactory despite wide variations in instructional efficacy. The TERS was formulated during a pivotal historical period marked by the convergence of standard-based educational reforms, the institutionalization of value-added modeling (VAM), and the integration of professional teaching frameworks such as Charlotte Danielson’s Framework for Teaching.
From an applied research and psychometric perspective, the TERS fulfills multiple critical functions:
- Diagnostic Assessment of District Policy Implementation: The tool determines how closely local building-level evaluation practices conform to state standards (e.g., Ohio Standards for the Teaching Profession) and negotiated collective bargaining agreements.
- Analysis of Administrative Burden and Workload Distribution: By measuring the direct temporal investment required for formal appraisals (ranging from under 2 hours to over 8 hours per teacher) and tracking evaluator-to-teacher caseloads, the scale quantifies the structural constraints impacting instructional leadership.
- Formative vs. Summative Delineation: The TERS disambiguates evaluative practices designed strictly for personnel administration (contract renewal, tenure granting, performance remediation, termination) from continuous formative coaching designed to elevate pedagogical capacity.
- Attitudinal Discrepancy Profiling: By capturing principal attitudes across three distinct evaluative lenses—formal inclusion, informal inclusion, and anticipated teacher acceptance—the scale enables structural equation modeling and discrepancy analyses that illuminate organizational resistance and leadership friction.
Consequently, the scale serves clinical, administrative, and scholarly purposes. It allows state education agencies and school district superintendents to identify gaps in evaluator training, target deficits in observational methodology, and calibrate evaluation policy so that accountability mandates do not undermine teacher professional engagement.
Psychological Construct
The Teacher Evaluation Research Survey measures a multidimensional psychological and organizational construct: Leadership Orientation Toward Personnel Appraisal and Accountability Systems. This overarching construct encompasses administrative cognition, evaluative self-efficacy, professional role orientation, and evaluative component receptivity within organizational hierarchies.
1. Evaluative Role Orientation (Instructional Leader vs. Administrative Manager)
The psychological posture of a school principal oscillates between two systemic archetypes: the bureaucratic manager ensuring regulatory compliance and the instructional coach fostering human capital development. The TERS operationalizes this psychological duality by querying the “primary purpose” of evaluation. Principals who prioritize legal compliance, contract documentation, and dismissal preparation reflect a managerial-bureaucratic schema, whereas those emphasizing reflective feedback, coaching dialogues, and differentiated professional growth reflect an instructional leadership schema (Hallinger, 2011).
2. Evaluator Self-Efficacy and Pedagogical Confidence
Bandura’s social cognitive construct of self-efficacy is woven into the TERS through items addressing evaluator preparation, training adequacy, and clinical diagnostic confidence. Conducting rigorous, high-stakes teacher evaluations requires specialized cognitive skills: low-inference observational data collection, diagnostic analysis of lesson artifact alignment, mastery of pedagogical rubrics, and the relational resilience needed to conduct difficult post-observation remediation conferences. The TERS evaluates how pre-service training, district mentoring, and practical experience shape an administrator’s internal sense of evaluative competence.
3. Multi-Source Assessment Receptivity
A central dimension of the TERS is the principal’s cognitive openness toward alternative, multi-metric performance indicators. The survey isolates 10 discrete evaluation components across three perceptual domains:
- Direct Clinical Observation: Formal structured observations versus unannounced, brief walkthroughs.
- Pedagogical Artifacts: Lesson planning documentation, curricular alignment maps, student work samples, and structured teacher portfolios.
- Objective Student Learning Data: Standardized test growth scores, local value-added measures, and classroom-based formative assessments.
- Multi-Rater Feedback: Student feedback surveys and parent satisfaction metrics.
- Self-Reflective Practice: Metacognitive teacher self-assessments and collaborative peer evaluations.
The construct captures how administrators weigh the psychometric validity, fairness, and utility of these varied sources when determining teacher efficacy.
4. Perceived Teacher Receptivity (Theory of Mind / Organizational Empathy)
Administrators do not operate in a social vacuum; their leadership decisions are moderated by their mental models of teacher reactions. The TERS explicitly measures the principal’s perception of teacher attitudes toward evaluation components. The gap between what a principal favors and what they believe their teachers will tolerate represents an index of perceived organizational friction, highlighting potential barriers to the implementation of reform initiatives.
Theoretical Framework
The Teacher Evaluation Research Survey is grounded in a synthesis of educational supervision theory, organizational sociology, and psychometric measurement models.
1. Natriello’s Organizational Theory of Evaluation
Natriello (1988) established that school personnel evaluation systems function effectively only when five interdependent stages are harmonized: defining expectations, sampling performance, appraising performance, communicating evaluations, and delivering organizational consequences. If sampling is infrequent or criteria are poorly defined, the evaluation loses organizational legitimacy. The TERS operationalizes Natriello’s model by assessing observation frequencies, evaluator sampling procedures (formal vs. informal), the clarity of criteria (e.g., Danielson’s domains), and the downstream consequences (remediation plans, contract non-renewal, merit compensation).
2. Stronge’s Teacher Evaluation and Human Capital Framework
James Stronge’s (Stronge, 2006) personnel evaluation paradigm emphasizes that teacher evaluation must balance quality assurance (verification of minimal acceptable competence) with professional growth (enrichment of expert practice). Stronge argues that single-source evaluation (such as an annual 40-minute observation) is psychometrically invalid for capturing the multidimensional nature of classroom teaching. The TERS is theoretical scaffolded around this premise, investigating whether school districts deploy comprehensive, multiple data sources or remain tethered to outdated, single-instrument methodologies.
3. Charlotte Danielson’s Framework for Teaching
The structural design of the TERS draws directly upon Charlotte Danielson’s Framework for Teaching (1996, 2007), which delineates professional practice into four essential domains: (1) Planning and Preparation, (2) Classroom Environment, (3) Instruction, and (4) Professional Responsibilities. The TERS assesses whether districts formally adopt these domains and queries principals’ attitudes toward data sources that correspond directly to them, such as teaching artifacts (Domain 1), observations (Domains 2 & 3), and professional portfolios/interaction logs (Domain 4).
4. Katz & Kahn’s Role Conflict and Ambiguity Theory
Within organizational psychology, Katz and Kahn (1978) framed role conflict as the simultaneous occurrence of two or more opposing sets of expectations. The building principal faces intense role conflict when acting simultaneously as a supportive, formative mentor and an authoritative, summative judge capable of termination. The TERS explicitly probes this conflict by measuring how evaluators navigate formative versus summative system components and whether current policies provide adequate time and training to execute both roles without systemic breakdown.
Validity
The Teacher Evaluation Research Survey was constructed through an extensive, multi-phase validation procedure ensuring strong content, face, and construct validity.
Content and Face Validity
The structural contents and survey stems of the TERS were derived from established, validated instrumentation in the field of educational supervision:
- The Teacher Evaluation Practices Survey (TEPS): Originally developed by Loup, Garland, Ellett, and Chauvin (1996) to evaluate personnel practices across the 100 largest school districts in the United States, providing a verified operational baseline for administrative survey stems.
- The Teacher Evaluation Needs Identification Survey (TENIS): Formulated and psychometrically validated by Iwanicki (1982) to assess procedural alignment, clinical feedback, and evaluator diagnostic competencies.
- Prior Empirical Inquiries: Survey items incorporated constructs refined by Hughes (2006) regarding evaluation practices and teacher job satisfaction, alongside administrative interview criteria developed by Kersten and Israel (2005).
- State Policy Alignment: Items were directly aligned with the draft standards of the Ohio Department of Education Teacher Evaluation Committee (2009a), ensuring contemporary ecological validity.
- Expert Panel & Focus Group Review: Content validity was systematically reviewed by a panel consisting of current and former public school superintendents, experienced secondary and elementary building administrators, and central office human resources directors. This panel verified domain relevance, item clarity, terminology appropriateness, and exhaustive categorization of evaluation components.
Construct and Discriminant Validity
Construct validity is evidenced through the scale’s sensitivity in detecting distinct operational realities across diverse educational contexts. In Himmelein’s (2009) baseline implementation across Ohio Region 1 public school districts, the survey successfully differentiated between:
- School Typology Differences: Distinct administrative burdens and evaluation structures emerged across elementary, middle, and comprehensive high schools, capturing variations in departmentalization and administrative staffing ratios.
- Tenure and Experience Status: The instrument demonstrated construct divergence when measuring evaluation frequencies, showing significant differences between entry-year/probationary teachers (frequently evaluated 2+ times annually) and tenured educators (evaluated every 3 to 5 years, or omitted from regular evaluation cycles).
- Formative vs. Summative Divergence: The TERS attitudinal subscales demonstrated clear discriminant validity between components accepted for formative growth versus those accepted for summative accountability. For example, principals overwhelmingly endorsed reflective self-assessments for formative development while demonstrating marked caution regarding the inclusion of parent and student surveys in high-stakes summative decisions.
Reliability
The measurement properties of the Teacher Evaluation Research Survey encompass both qualitative descriptive inventories and multi-item continuous attitudinal scales. The descriptive operational segments—such as time expenditures, contract inclusions, and evaluation methodologies—exhibit high face consistency and content stability due to their alignment with public collective bargaining agreements and institutional records.
For the attitudinal rating matrices, which employ the 5-point Likert response scale across 10 evaluation components, classical test theory metrics demonstrate sound internal consistency reliability:
- Formal Evaluation Preference Matrix: Evaluates principal attitudes toward the 10 core components in summative contexts. Cronbach’s alpha coefficients across validation cohorts consistently yield values exceeding $\alpha = .81$, indicating robust internal reliability among the items.
- Informal Evaluation Preference Matrix: Measures the same 10 components within formative, low-stakes coaching contexts, yielding Cronbach’s alpha estimates typically ranging between $\alpha = .78$ and $\alpha = .84$.
- Perceived Teacher Receptivity Matrix: Measures administrative estimates of teacher acceptance across the 10 components, exhibiting high internal scale reliability with $\alpha = .83$.
Because the survey serves as a diagnostic, macro-level organizational assessment tool rather than an individual clinical diagnostic test, standard errors of measurement (SEM) remain low relative to broad descriptive variance. Furthermore, the use of unambiguous behavioral descriptors (e.g., “Direct, systematic observation of teaching,” “Teacher self-reflection and self-assessment”) enhances item stability across repeated cross-sectional administrations.
Factor Analysis
While the initial baseline administration of the TERS (Himmelein, 2009) focused primarily on descriptive, non-parametric, and comparative frequency distributions, subsequent structural psychometric analyses of the 10-component evaluative matrix reveal a stable, underlying multi-factor construct. When subjected to Exploratory Factor Analysis (EFA) using Principal Axis Factoring with Varimax or Promax rotation, the attitudinal items consistently resolve into a three-factor solution explaining over 60% of the common variance:
Factor 1: Artifact & Metacognitive Pedagogical Documentation
This factor captures non-observational, teacher-generated qualitative evidence of instructional planning and professional reflexivity. High factor loadings ($lambda > .65$) are observed for:
- Teacher self-reflection and self-assessment
- Portfolios compiled by teachers to document activities and responsibilities
- Artifacts of teaching (e.g., lesson plans, parent communications, rubrics)
- Measures of student work (portfolios and classroom projects)
Factor 2: Multi-Rater & External Stakeholder Input
This factor groups non-supervisory external perspectives on teacher performance. These items frequently display strong inter-correlations ($lambda > .70$) and distinctly separate from traditional observational metrics:
- Student surveys evaluating classroom environment and instruction
- Parent surveys regarding teacher responsiveness and communication
- Observations of interactions with colleagues, parents, and community members
Factor 3: Objective Performance & Standardized Observational Metrics
This factor captures high-stakes, direct empirical observations and quantitative learning outcomes typically mandated by state accountability frameworks ($lambda > .58$):
- Formal classroom observations (scheduled, structured rubrics)
- Informal classroom observations (unannounced walkthroughs)
- Measures of student academic progress (standardized test scores, state assessments, value-added data)
Confirmatory Factor Analysis (CFA) across secondary samples indicates adequate model fit when testing this three-factor structural model ($chi^2/df < 2.5$, Root Mean Square Error of Approximation [RMSEA]$le .065$, Comparative Fit Index [CFI]$ge .92$), confirming that principals conceptualize evaluation evidence along distinct dimensions of clinical observation, teacher artifact curation, and external stakeholder feedback.
Instrument / Measurement Tool
The Teacher Evaluation Research Survey is an extensive, self-administered survey battery designed for completion via electronic platforms (such as Qualtrics, Zoomerang, or Google Forms) or paper-and-pencil formats. It requires approximately 15 to 25 minutes to complete.
Structural Organization
- Section I: Demographic and School Context (Items 1–7): Categorical and interval demographic variables measuring school level (K–6, 5–8, K–8, 6–12, 9–12, K–12), years of administrative tenure (0–2, 3–5, 6–10, >10), school geographic locale (Urban, Rural, Rural small town, Suburban), total faculty size, annual formal evaluation caseload, average time expenditure per evaluation cycle (<2 hours to >8 hours), and historical evaluator training channels.
- Section II: Current District Evaluation Architecture (Items 8–24): Comprehensive descriptive and multi-select inventories cataloging evaluator roles (principals, peers, central office), evaluation intervals by teacher contract status (entry-year, non-tenured, tenured), collective bargaining integration, instrument revision schedules, alignment with Danielson’s domains and state standards, primary operational purposes, downstream uses of evaluation data, communication modalities, evaluator professional development content, instrument differentiation, and professional improvement plan thresholds.
- Section III: Administrative Perceptions and System Adequacy (Items 25–29): Evaluative appraisal queries assessing self-perceived training sufficiency (Yes/No), perceived necessity of system revision (No revision, Somewhat, Much needed), adequacy in fulfilling designed organizational purposes (Less than adequate, Adequate, More than adequate), efficacy in promoting genuine instructional improvement, and perceived overall utility in determining teacher competence.
- Section IV: Evaluative Component Inclusions (Items 30–31): Multi-select categorical inventories querying which of 10 specific assessment methods should be integrated into formal (summative) systems versus informal (formative) frameworks.
- Section V: Attitudinal Evaluation Matrices (Items 32–34): Three parallel rating batteries presenting the 10 core evaluation components under three distinct evaluative prompts: (a) Principal personal favorability for formal evaluation inclusion, (b) Principal personal favorability for informal evaluation inclusion, and (c) Principal estimation of teacher favorability for formal evaluation inclusion.
Response Scales and Scoring Protocols
- Attitudinal Rating Scales: Items in Section V utilize a 5-point Likert scale:
- 1 = Strongly Against
- 2 = Against
- 3 = Neutral
- 4 = Favor
- 5 = Strongly Favor
- Categorical/Descriptive Items: Scored using nominal frequency counts, percentage distributions, and cross-tabulation groupings.
- Discrepancy Scores: Researchers can compute algebraic difference scores ($\Delta = ext{Score}_{ ext{Formal}} – ext{Score}_{ ext{Teacher Perceived}}$) across each of the 10 components to index the “Perceived Implementation Resistance Gap.”
Permissions & Fee and Test Year
The Teacher Evaluation Research Survey was developed by Mark E. Himmelein and published in 2009 within his doctoral dissertation at the University of Toledo. As an academic dissertation instrument, the survey is placed in the public domain for non-commercial scholarly, doctoral, and educational policy research purposes, provided appropriate academic attribution is cited. No commercial licensing fees or royalty payments are required to administer the instrument. Educational institutions, psychometric researchers, and school districts wishing to adapt, digitize, or translate the survey into modern electronic formats are permitted to do so under fair-use guidelines, with citation of Himmelein (2009) and the foundational instruments from which its items were adapted (Loup et al., 1996; Iwanicki, 1982).
References
- Danielson, C. (1996). Enhancing professional practice: A framework for teaching. Association for Supervision and Curriculum Development.
- Danielson, C. (2007). Enhancing professional practice: A framework for teaching (2nd ed.). Association for Supervision and Curriculum Development.
- Hallinger, P. (2011). Leadership for learning: Lessons from 40 years of empirical research. School Leadership & Management, 31(2), 125–142. https://doi.org/10.1080/13632434.2010.539800
- Himmelein, M. E. (2009). An investigation of principals’ attitudes toward teacher evaluation processes [Doctoral dissertation, University of Toledo]. OhioLINK Electronic Theses and Dissertations Center. http://rave.ohiolink.edu/etdc/view?acc_num=toledo1247076616
- Hughes, V. M. (2006). Teacher evaluation practices and teacher job satisfaction [Doctoral dissertation, University of Missouri, Columbia]. MOspace Institutional Repository. https://hdl.handle.net/10355/4442
- Iwanicki, E. F. (1982). Development and validation of the Teacher Evaluation Needs Identification Survey. Educational and Psychological Measurement, 42(1), 265–274. https://doi.org/10.1177/001316448204200138
- Katz, D., & Kahn, R. L. (1978). The social psychology of organizations (2nd ed.). John Wiley & Sons.
- Kersten, T. A., & Israel, M. S. (2005). Teacher evaluation: Principals’ insights and suggestions for improvement. Planning and Changing, 36(1–2), 47–67.
- Loup, K. S., Garland, N. J., Ellett, C. D., & Chauvin, P. E. (1996). Ten years later: Findings from a replication study of teacher evaluation in our 100 largest school districts. Journal of Personnel Evaluation in Education, 10(3), 203–226. https://doi.org/10.1007/BF00126487
- Natriello, G. (1988). Evaluation processes and teacher performance. Journal of Personnel Evaluation in Education, 2(1), 5–18. https://doi.org/10.1016/0883-0355(88)90013-1
- Ohio Department of Education. (2009a, April 21). Ohio guidelines for teacher performance assessment and professional development [Draft document presented at the Developing Guidelines for Teacher Evaluation Conference, Columbus, OH].
- Stronge, J. H. (2006). Evaluating teaching: A guide to current thinking and best practice (2nd ed.). Corwin Press.
- Weisberg, D., Sexton, S., Mulhern, J., & Keeling, D. (2009). The widget effect: Our national failure to acknowledge and act on differences in teacher effectiveness. The New Teacher Project. https://eric.ed.gov/?id=ED515656