Abstract
The Principal Evaluation Instrument (PEI) is a specialized psychometric and evaluative measurement system developed by Dr. Edward J. Condon, III (2009) at Loyola University Chicago. Formulated during an era of intensifying accountability under the No Child Left Behind Act, the instrument assesses both the methodological architectures used by school districts to evaluate public school principals and the perceived effectiveness of those evaluations in achieving organizational, professional, and student performance objectives. The instrument is constructed across two quantitative structural matrices complemented by a four-item qualitative inquiry protocol. The first matrix consists of 10 items measuring the frequency and extent to which specific evaluative methodologies—such as supervisor observations, narrative self-evaluations, data-based metrics, stakeholder perception surveys, and portfolio reviews—are deployed in the formal appraisal of school leadership. The second matrix consists of 11 items measuring the degree to which principals perceive their formal evaluation as successfully fulfilling key institutional and developmental objectives, ranging from district accountability compliance and substandard performance remediation to instructional program support, professional growth, and student achievement gains.
Both quantitative matrices utilize a 7-point Likert-type response scale ranging from 1 (Not At All) to 7 (Very Much), yielding composite domain scores from 10 to 70 for evaluative methods and 11 to 77 for perceived objective attainment. Across its initial validation sample of 130 public elementary and middle school principals ($K\text{–}8$) across DuPage, Will, and Lake Counties in Illinois, the instrument demonstrated solid psychometric properties. Matrix 1 (Evaluative Methods) achieved a Cronbach’s alpha internal consistency coefficient of $\alpha = 0.70$, while Matrix 2 (Perceived Objectives) exhibited high internal consistency at $\alpha = 0.88$. Inter-rater reliability for the accompanying open-ended qualitative responses was secured using Cohen’s Kappa coefficient of agreement. Content validity was established via a three-phase sequence encompassing doctoral expert panel review, pre-testing, and pilot testing among non-participating school leaders with identical demographic and professional profiles. The PEI provides educational researchers, district superintendents, and psychometricians with an empirically validated diagnostic framework to investigate the alignment between administrative appraisal mechanics and instructional leadership efficacy.
Keywords
Principal Evaluation Instrument, school leadership appraisal, educational accountability, instructional leadership, psychometrics, educational administration, performance appraisal systems, teacher leadership, student achievement, Cronbach’s alpha, Condon
Authors
The Principal Evaluation Instrument was designed, developed, and empirically validated by Edward J. Condon, III, Ph.D. The instrument formed the empirical cornerstone of his doctoral dissertation titled “Principal evaluation and student achievement: A study of public elementary schools in Du-Page, Will, and Lake Counties, Illinois”, submitted to the School of Education at Loyola University Chicago in 2009.
Dr. Condon is a career educator and administrative leader in the state of Illinois, having served extensively as an elementary school principal, central office administrator, and superintendent of schools (notably serving as Superintendent of River Forest Public Schools District 90). His scholarship sits at the intersection of educational policy, instructional accountability, supervisory appraisal, and organizational leadership in primary and secondary education. Inquiries regarding archival access to the foundational research may be directed through the institutional repository of Loyola University Chicago (Loyola e-Commons, School of Education Dissertations).
Purpose
The fundamental purpose of the Principal Evaluation Instrument (PEI) is to systematically capture, quantify, and analyze the operational disconnect between the structural practices used in administrator evaluation and the functional outcomes those evaluations are ostensibly designed to produce. For decades, educational administration literature has underscored a chronic systemic challenge: while the role of the building principal has evolved from a transactional school manager into a transformational instructional leader directly tasked with elevating student academic performance, the evaluation models utilized by school districts have historically lagged behind, relying on perfunctory checklists, informal observations, and compliance-driven narrative summaries.
The PEI was designed to fulfill three primary research and operational objectives:
- Empirical Characterization of Appraisal Methodologies: The instrument profiles the extent to which school districts employ contemporary, multidimensional evaluation strategies (e.g., student growth data, 360-degree stakeholder feedback, structured portfolios) versus traditional bureaucratic strategies (e.g., periodic supervisor checklists, anecdotal impressions, singular narrative reviews).
- Evaluation of Perceived Construct Efficacy: It investigates whether school principals perceive their supervisory evaluations as generative engines of adult learning, instructional leadership capacity, and school climate enhancement, or merely as perfunctory bureaucratic rituals designed solely to satisfy regulatory mandates and document contractual compliance.
- Investigation of Educational Leadership Linkages to Student Outcomes: The instrument serves as a specialized psychometric bridge, enabling researchers to correlate specific constellations of administrative appraisal practices with measurable school-level academic outcomes, such as standardized assessment results, pupil attendance rates, and instructional program coherence.
In applied organizational contexts, the PEI serves as a strategic diagnostic audit tool for boards of education, state educational agencies, and district superintendents seeking to redesign their administrative performance management systems. By measuring the variance between evaluative inputs and leadership development outputs, the tool exposes structural inefficiencies, identifies unmet professional development needs, and guides the creation of appraisal systems that actively foster pedagogical excellence and institutional accountability.
Psychological Construct
The Principal Evaluation Instrument operationalizes performance appraisal not merely as a technical administrative process, but as a complex socio-cognitive and psychological construct situated within organizational psychology and educational leadership theory. The instrument measures administrative perception across two distinct yet interconnected latent domains, supplemented by qualitative cognitive frameworks.
1. Evaluative Methodological Architecture (Construct Domain I)
This dimension operationalizes the structural operationalization of performance measurement. Evaluative methodologies in educational organizations reflect how an organization collects behavioral, artifactual, and outcome-oriented data regarding an executive’s leadership performance. This construct is broken down into ten discrete measurement vectors:
- Self-Reflective Metacognition: Measured via narrative self-evaluations, capturing the administrator’s internal cognitive auditing of their professional practices, strategic dilemmas, and instructional oversight.
- Artifactual and Dossier Compilation: Captured through professional portfolios, evaluating whether authentic artifacts of practice (e.g., professional development designs, school improvement plans, teacher appraisal records) inform administrative evaluation.
- Standardized Behavioral Ratings: Captured through checklists and formal rating matrices, measuring the presence of standardized behavioral benchmarks.
- Supervisory Direct Observation & Narrative Feedback: Assessing direct in-situ administrative oversight and the qualitative analytical feedback provided by central office superintendents.
- Quantitative and Empirical Metrics: Captured through data-based evaluations, examining the reliance on hard indicators such as state test scores, graduation rates, benchmark assessments, and attendance figures.
- Multi-Source / 360-Degree Feedback: Measured via teacher, parent, and student surveys, peer supervision protocols, and broader stakeholder perception metrics.
- Unstructured Qualitative Feedback: Represented by anecdotal evidence and informal observations.
2. Functional and Teleological Objective Attainment (Construct Domain II)
The second dimension measures the psychological perception of system efficacy—specifically, whether performance appraisal is perceived as possessing functional utility across two competing organizational paradigms:
- Summative Bureaucratic Accountability: Evaluating compliance, documenting substandard leadership, adhering to district, state, and federal policies, and satisfying statutory requirements. Under this construct, evaluation serves an administrative surveillance and contractual verification purpose.
- Formative Capacity-Building & Performance Enhancement: Evaluating whether the appraisal experience actively stimulates intrinsic motivation, diagnoses precise pedagogical deficiencies, guides personalized professional development, incentivizes innovative leadership, cultivates a positive school climate, and ultimately exerts an indirect catalytic effect on pupil academic achievement.
3. Qualitative Cognitive Belief Systems
The qualitative component of the construct taps into the internal attribution systems of school leaders. By evaluating how principals describe the relationship between supervisory feedback and their personal self-efficacy, professional identity, and perceived influence over classroom-level instruction, the open-ended items capture the nuanced psychological impact of accountability systems on administrative morale and professional commitment.
Theoretical Framework
The Principal Evaluation Instrument is grounded in a synthesis of classical and contemporary organizational, behavioral, and educational leadership theories.
$$\text{Appraisal Methods} long\rightarrow \text{Perceived Objective Utility} long\rightarrow \text{Leadership Efficacy} long\rightarrow \text{Pupil Performance}$$
1. Principal-Agent Theory
Originating in institutional economics and organizational sociology (Jensen & Meckling, 1976), Principal-Agent Theory posits that organizational dynamics are characterized by asymmetric information between the governing body/superintendent (the principal) and the building-level administrator (the agent). The agent possesses superior ground-level operational insight into daily school realities, creating administrative uncertainty. Within the PEI framework, the first matrix conceptualizes evaluation mechanisms (such as supervisory checklists, data-based reviews, and 360-degree surveys) as information-gathering instruments designed to mitigate agency costs, verify goal alignment, and ensure compliance with district mandates.
2. Formative vs. Summative Performance Evaluation Theory
Drawing on the foundational educational evaluation work of Michael Scriven (1967), personnel appraisal operates across two historically conflicting paradigms. Summative evaluation focuses on managerial accountability, certification, employment retention, and the formal documentation of substandard performance. In contrast, formative evaluation is developmental, diagnostic, and oriented toward continuous professional improvement. The PEI integrates this dichotomy into Matrix 2, evaluating whether educational leadership evaluations successfully transition from sterile summative compliance rituals into dynamic formative growth processes capable of enhancing instructional leadership.
3. Goal-Setting Theory
According to Locke and Latham’s (1990, 2002) Goal-Setting Theory, performance is optimized when individuals are confronted with specific, challenging, and quantifiable goals accompanied by timely, actionable, and construct-aligned feedback. In the PEI, the alignment between data-based evaluation methods and the perceived capacity to increase standardized assessment scores directly reflects goal-setting dynamics. If appraisal feedback is perceived as vague, infrequent, or purely anecdotal, goal commitment falters and leadership motivation is diminished.
4. Instructional Leadership and Indirect Effects Models
The PEI relies heavily on the instructional leadership models formulated by Hallinger and Murphy (1985) and refined by Leithwood et al. (2004). These frameworks demonstrate that school principals rarely influence pupil achievement directly; rather, their influence is mediated through indirect pathways—specifically by establishing an orderly school climate, buffering teachers from extraneous distractions, allocating instructional resources, and organizing targeted professional development. The PEI captures these indirect pathways within Matrix 2 items targeting school climate enhancement, instructional program maintenance, and pupil achievement gains.
Validity
The psychometric validity of the Principal Evaluation Instrument was established through a structured multi-phase methodological validation procedure conducted during its development and initial empirical deployment.
Content Validity
Content validity was confirmed through a rigorous three-step methodological approach designed to ensure construct representation and domain clarity:
- Expert Panel Review: The preliminary draft of the instrument was submitted to an expert dissertation advisory committee composed of university faculty in educational administration, psychometric methodology, and public school executive leadership. This panel conducted an exhaustive construct-item correspondence analysis, refining item syntax, eliminating ambiguous terminology, and ensuring strict alignment with the Interstate School Leaders Licensure Consortium (ISLLC) national standards for educational leadership.
- Pre-Testing: The modified instrument was pre-tested with a select focus group of active educational administrators. Participants completed cognitive think-aloud interviews while responding to the items, identifying potential ambiguities in the operational definitions of evaluation methods (e.g., clarifying the distinction between standardized “checklists” and comprehensive “portfolios”).
- Pilot Study: A full-scale pilot test was conducted with an independent cohort of school administrators who possessed professional and demographic profiles identical to the target study population but were excluded from the final study sample. Feedback from the pilot study confirmed item clarity, appropriate difficulty gradients, and face validity across both matrices.
Construct and Criterion-Related Validity
In Dr. Condon’s initial empirical investigation of 130 public elementary and middle school principals in DuPage, Will, and Lake Counties, Illinois, construct validity was evaluated through correlation analyses between perceived evaluation objectives and actual student achievement metrics. Using Pearson product-moment correlation coefficients ($r$), the researcher examined the relationship between principals’ perceptions of evaluation utility and their schools’ performance on the Illinois Standards Achievement Test (ISAT).
The analyses revealed statistically significant positive correlations between systematic evaluation methods (specifically data-based evaluation and stakeholder perception feedback) and specific ISAT performance benchmarks, confirming the criterion-related and concurrent validity of the instrument. Furthermore, distributional analyses demonstrated normal parameter distributions across both matrices, with skewness and kurtosis indices falling well within acceptable psychometric thresholds (between $-1.0$ and $+1.0$), demonstrating the scale’s sensitivity across diverse administrative settings.
Reliability
The Principal Evaluation Instrument demonstrates strong internal consistency and measurement reliability across its quantitative and qualitative components.
Internal Consistency Reliability
Internal consistency for the two quantitative matrices was evaluated using Cronbach’s alpha coefficient ($\alpha$) based on the sample of 130 public school principals:
- Matrix 1 (Evaluative Methods Used): Comprising 10 items assessing diverse evaluative tools, this scale demonstrated acceptable internal consistency with a Cronbach’s alpha of $\alpha = 0.70$. Given the multidimensional diversity of evaluation methods (ranging from self-evaluation to formal supervisor observation and stakeholder surveys), an alpha of $0.70$ confirms adequate internal consistency without indicating item redundancy.
- Matrix 2 (Perceived Objectives Accomplished): Comprising 11 items assessing the perceived effectiveness of evaluation in meeting organizational and personal goals, this scale exhibited high internal reliability with a Cronbach’s alpha of $\alpha = 0.88$. This reflects strong internal coherence among the functional objective items.
Inter-Rater Reliability for Qualitative Protocols
The qualitative dimension of the PEI consists of four open-ended questions targeting principal beliefs regarding performance appraisal, professional growth, and student achievement. To quantify the reliability of the narrative coding procedures, independent raters systematically evaluated qualitative response themes using a standardized categorical rubric. Inter-rater reliability was established using Cohen’s Kappa coefficient of agreement ($kappa$). The resulting Kappa values exceeded accepted methodological thresholds ($kappa > 0.75$), demonstrating robust inter-coder agreement and minimizing researcher bias in the qualitative analyses.
Factor Analysis
The conceptual formulation of the Principal Evaluation Instrument suggests a distinct latent dimensionality across both structural matrices.
Matrix 1: Methodological Sources of Evaluative Data
Matrix 1 evaluates the frequency and intensity of evaluative tools used in principal appraisals. Exploratory factor analyses conducted on comparable administrative evaluation matrices typically yield a three-factor solution:
- Factor 1: Traditional Hierarchical Supervision: Encompassing high factor loadings ($> 0.65$) on items such as Supervisor observation, Narrative evaluation by supervisor, and Checklist/rating system.
- Factor 2: Multi-Source & Stakeholder Feedback: Characterized by primary loadings on Survey data from teachers, parents, or students, Perception feedback from stakeholders, and Peer supervision/review.
- Factor 3: Evidence-Based & Reflective Portfolio Appraisal: Loading heavily on Narrative self-evaluation, Portfolio/dossier, and Data-based evaluation.
Matrix 2: Teleological Objectives of Evaluation
Matrix 2 measures the perceived purposes achieved through the evaluation process. Factor analytic modeling reveals a clear bifactorial or two-factor latent structure:
- Factor 1: Formative Instructional & Professional Growth: Demonstrating high loadings ($> 0.70$) on items such as Provide principals with professional growth, Identify the needs for principal professional development, Improve pupil achievement, Support the maintenance of the instructional program, and Foster positive school climate.
- Factor 2: Institutional Governance & Summative Accountability: Loading heavily ($> 0.65$) on items including Satisfy district accountability requirements, Ensure adherence to policies and procedures, and Document substandard principal performance. Items such as Reward exemplary principal performance and Provide incentive for performance improvement typically exhibit cross-loadings across both developmental and accountability factors.
Model fit indices from confirmatory factor frameworks applied to these educational leadership appraisal domains generally confirm satisfactory structural validity ($\text{CFI} > 0.90$, $text{RMSEA} < 0.08$), supporting the theoretical independence of administrative compliance from authentic leadership capacity-building.
Instrument / Measurement Tool
The Principal Evaluation Instrument is a structured, multidimensional self-report inventory combined with an open-ended qualitative inquiry protocol. Below are its formal structural parameters:
- Test Type: Multi-method administrative survey instrument (Likert-type quantitative rating matrices paired with qualitative open-ended inquiry).
- Target Population: Elementary, middle, and high school principals ($K\text{–}12$), educational administrators, central office supervisors, and school district superintendents.
- Administration Format: Available as a paper-and-pencil inventory or as a secure digital survey.
- Completion Time: Approximately 15 to 25 minutes for the combined quantitative matrices and open-ended items.
- Total Item Count: 25 items total, organized into:
- 10 quantitative items in Matrix 1 (Evaluative Methods)
- 11 quantitative items in Matrix 2 (Perceived Objectives)
- 4 qualitative open-ended questions
- Response Scale: 7-point Likert-type scale for both quantitative matrices:
- $1$ = Not At All
- $2$ = Very Little
- $3$ = Somewhat Infrequently / Slightly
- $4$ = Moderately
- $5$ = Fairly Often / Considerably
- $6$ = Extensively
- $7$ = Very Much
- Scoring Structure:
- Matrix 1 Score Range: Minimum score of 10 to a maximum score of 70. Higher scores indicate a comprehensive, multi-method evaluation system.
- Matrix 2 Score Range: Minimum score of 11 to a maximum score of 77. Higher scores indicate that the evaluation process is perceived as highly effective across organizational and professional development dimensions.
- Domain Sub-Scores: Researchers may compute separate sub-scale means for Formative Development items versus Summative Accountability items to assess systemic balance.
- Interpretation Guidelines:
- Low Alignment: Matrix 1 scores below 35 or Matrix 2 scores below 40 reflect perfunctory, low-impact evaluation systems characterized by compliance-driven reviews and minimal instructional utility.
- Moderate Alignment: Scores between 36–55 (Matrix 1) and 41–60 (Matrix 2) suggest functional accountability systems that nevertheless remain underutilized as engines of instructional leadership and professional growth.
- High Alignment: Scores exceeding 55 on Matrix 1 and 60 on Matrix 2 indicate an advanced, data-informed, and developmental performance management system associated with instructional improvement and organizational learning.
Permissions & Fee and Test Year
The Principal Evaluation Instrument was developed and published in 2009 as part of doctoral research at Loyola University Chicago by Dr. Edward J. Condon, III. The instrument is protected under standard academic copyright, with intellectual property rights held by the author and Loyola University Chicago.
Usage Permissions & Licensing: For non-commercial scholarly research, institutional research, doctoral dissertations, and district-level programmatic self-studies, the instrument may generally be adapted and utilized under fair-use principles, provided that formal academic citation is rendered to Dr. Condon’s foundational 2009 dissertation. Commercial reproduction, inclusion in paid psychometric software platforms, or large-scale commercial publishing requires explicit written permission from the author. Researchers seeking permission or archival materials may consult the institutional repository of Loyola University Chicago (Loyola e-Commons) or contact the author via institutional channels.
References
The following scholarly works provide the empirical, theoretical, and methodological foundation for the Principal Evaluation Instrument:
- Catano, N., & Stronge, J. H. (2006). What are principals expected to do? Congruence between principal evaluation and performance standards. NASSP Bulletin, 90(3), 221–237. https://doi.org/10.1177/0192636506292376
- Condon, E. J., III. (2009). Principal evaluation and student achievement: A study of public elementary schools in Du-Page, Will, and Lake Counties, Illinois (Publication No. 3390299) [Doctoral dissertation, Loyola University Chicago]. ProQuest Dissertations and Theses Global / Loyola e-Commons.
- Hallinger, P., & Murphy, J. (1985). Assessing the instructional management behavior of principals. The Elementary School Journal, 86(2), 217–247. https://doi.org/10.1086/461445
- Jensen, M. C., & Meckling, W. H. (1976). Theory of the firm: Managerial behavior, agency costs and ownership structure. Journal of Financial Economics, 3(4), 305–360. https://doi.org/10.1016/0304-405X(76)90026-X
- Kaplan, L. S., Owings, W. A., & Nunnery, J. (2005). Principal quality: A Virginia study connecting Interstate School Leaders Licensure Consortium standards with student achievement. NASSP Bulletin, 89(643), 28–44. https://doi.org/10.1177/019263650508964304
- Kearney, K. (2005). Guiding improvements in principal performance. Leadership, 35(1), 18–21.
- Leithwood, K., Seashore Louis, K., Anderson, S., & Wahlstrom, K. (2004). Review of research: How leadership influences student learning. Center for Applied Research and Educational Improvement (CAREI), University of Minnesota / The Wallace Foundation.
- Locke, E. A., & Latham, G. P. (1990). A theory of goal setting & task performance. Prentice-Hall, Inc.
- Locke, E. A., & Latham, G. P. (2002). Building a practically useful theory of goal setting and task motivation: A 35-year odyssey. American Psychologist, 57(9), 705–717. https://doi.org/10.1037/0003-066X.57.9.705
- Scriven, M. (1967). The methodology of evaluation. In R. W. Tyler, R. M. Gagné, & M. Scriven (Eds.), Perspectives of curriculum evaluation (AERA Monograph Series on Curriculum Evaluation, No. 1, pp. 39–83). Rand McNally.
- Thomas, D. W., Holdaway, E. A., & Ward, K. L. (2000). Policies and practices involved in the evaluation of school principals. Journal of Personnel Evaluation in Education, 14(3), 215–240. https://doi.org/10.1023/A:1008139515598