Abstract
The Münster Questionnaire for the Evaluation of Exams (German: Münsteraner Fragebogen zur Evaluation von Klausuren, abbreviated as MFE-K) is a specialized psychometric assessment instrument developed at the University of Münster (Westfälische Wilhelms-Universität Münster) to evaluate the quality, didactic alignment, administrative framing, and perceived student burden associated with written university examinations. Originally designed as an extension module within the broader institutional evaluation framework ("Münster Questionnaire for Evaluation"), the MFE-K addresses a critical systemic deficit in higher education quality assurance: while student evaluations of teaching (SET) routinely examine lectures and seminars, examinations themselves have historically lacked standardized, empirically validated diagnostic instruments. The core instrument captures multiple facets of the examination experience across 27 items, encompassing dimensions such as exam adequacy and construction quality, alignment with lecture content, environmental and administrative conditions, individual student exam preparation, and anticipated grades or global evaluative impressions. Responses are primarily recorded using a seven-point Likert-type agreement scale ranging from 1 (stimme gar nicht zu / strongly disagree) to 7 (stimme vollkommen zu / strongly agree), supplemented by specific open-ended and categorical metric indicators. Psychometric investigations reveal sound internal consistency across core subscales, with Cronbach’s alpha values typically ranging between .76 and .84. Confirmatory factor analyses support multi-dimensional models capturing construct validity, and multivariate analyses demonstrate high sensitivity to between-exam differences (η2 = .22). The MFE-K serves as an essential empirical feedback tool for university faculties, instructional development centers, and educational researchers striving to align higher education testing with modern competency-oriented assessment principles.
Keywords
Münster Questionnaire for the Evaluation of Exams, MFE-K, exam evaluation, higher education assessment, psychometrics, course evaluation, student burden, instructional quality, Klausurevaluation, academic testing quality
Authors
The Münster Questionnaire for the Evaluation of Exams was developed by educational and organizational psychologists affiliated with the Department of Psychology at the Westfälische Wilhelms-Universität (WWU) Münster, in institutional collaboration with evaluation specialists:
- Dr. Meinald T. Thielsch — Department of Psychology, Westfälische Wilhelms-Universität Münster, Fliednerstraße 21, 48149 Münster, Germany. Email: [email protected]. His research focuses on psychological diagnostics, usability, evaluation methodology, and organizational psychology.
- Benjamin Froncek, M.A. — FernUniversität in Hagen, Institute of Psychology, Universitätsstraße 33, 58097 Hagen, Germany. Email: [email protected]. Specializing in educational evaluation, academic achievement, and survey methodology.
- Collaborative Working Groups — Contributions were made by departmental evaluation commissions, including research teams at WWU Münster (e.g., Thielsch et al., 2008, 2010; Bechler & Thielsch, 2012) and student representative bodies (Fachschaft Psychologie), supporting ongoing item revisions and stakeholder-centered didactic refinements.
Purpose
The primary purpose of the MFE-K is to provide a standardized, psychometrically grounded methodology for measuring student perceptions of written summative examinations in higher education. Within modern university structures, written examinations ("Klausuren") serve as the predominant vehicle for evaluating student learning outcomes, awarding credits, and establishing academic progression criteria. Despite their central pedagogical and legal importance, examinations historically represented an unmonitored blind spot in academic quality management. While instructors frequently received feedback regarding their presentation style, course materials, or interpersonal rapport via traditional course evaluations, written exams were rarely subjected to systematic, psychometrically validated student evaluations.
This deficit became acute following the structural reorganization of European higher education under the Bologna Process. The replacement of cumulative diploma degree examinations with continuous, semester-concurrent modular testing precipitated a dramatic increase in examination density. As noted by higher education didactic researchers, students faced an "explosive proliferation" of exams, producing heightened cognitive load, strategic learning behaviors, and schedule compression. The German Standing Conference of the Ministers of Education and Cultural Affairs (Kultusministerkonferenz, KMK) explicitly highlighted an urgent mandate for action regarding the quality control, fairness, and transparency of university testing practices.
The MFE-K directly addresses this mandate across multiple clinical, administrative, and research domains:
- Didactic Quality Development: It delivers granular, actionable diagnostic feedback to examiners regarding question clarity, construct relevance, content representativeness, and structural transparency, allowing academic departments to identify flawed item formats, ambiguities, or unrealistic time limits.
- Monitoring Student Stress and Workload: By assessing the perceived cognitive and temporal strain associated with preparation and scheduling conflicts, faculty leadership can detect curriculum bottlenecks and prevent acute student burnout.
- Construct Alignment Verification: The instrument assesses whether the tested material mirrors the core curricula of lectures and whether assessments demand higher-order cognitive application versus superficial rote memorization.
- Legal and Organizational Fairness: By systematically tracking environmental conditions (room temperature, acoustic disruptions, lighting) and organizational execution (invigilation professionalism, point allocation clarity), the scale supports procedural fairness and institutional compliance.
Psychological Construct
The MFE-K evaluates the multi-dimensional construct of examination quality as experienced by examinees immediately post-assessment. Rather than viewing an exam solely through an objective classical test theory lens (e.g., item difficulty index p and point-biserial discrimination rit), the MFE-K captures the ecological validity, perceived fairness, and psychological burden of the testing episode. Across its comprehensive 27-item architecture, the instrument operationalizes five key psychological and didactic dimensions:
1. Exam Adequacy and Construction Quality (Angemessenheit und Güte der Klausur)
This core dimension measures the perceived psychometric and didactic craftsmanship of the examination. It addresses whether test questions were formulated unambiguously, whether difficulty levels were appropriately calibrated to the instructional level, whether point weightings were fair, and whether testing times were sufficient. Crucially, it captures whether the assessment evaluated profound conceptual comprehension rather than superficial mechanical regurgitation, as well as the overall construct representativeness of the question sample.
2. Curricular Alignment with Instructional Content (Übereinstimmung mit der Vorlesung)
Grounding itself in the educational principle of constructive alignment, this construct assesses the concordance between the taught curriculum and the tested curriculum. It examines whether the thematic priorities established during lectures were faithfully represented in the examination items, whether content surprised students unfairly, and whether the scope of tested materials strictly mirrored the authorized course syllabus.
3. Environmental and Administrative Framework (Rahmenbedingungen)
This dimension operationalizes the contextual conditions under which the test occurred. Standardized test administration requires that extraneous sources of measurement error be minimized. This subscale captures spatial factors (room acoustics, ergonomic space, ambient lighting, temperature control), invigilation decorum (maintenance of a quiet, uninterrupted setting, professional resolution of procedural queries), and the technical legibility of printed or digital exam interfaces.
4. Student Preparation and Learning Engagement (Vorbereitung)
Exam evaluations cannot be interpreted in a vacuum independent of student effort. This dimension measures the student’s self-regulated preparation intensity, total invested study hours, structured time management, and thoroughness of review. Capturing preparation allows researchers and evaluators to contextualize negative feedback: ratings from highly prepared students carry different didactic implications than those driven by procrastination or inadequate effort.
5. Grade Expectation and Global Evaluation (Notenerwartung und Gesamteindruck)
This dimension gathers summative global judgments and subjective cognitive projections. It captures expected academic performance, school-style global grading of the exam paper itself, and open-ended qualitative reflections (praise, substantive critique, didactic recommendations). By recording expected grades prior to actual grade dissemination, the instrument controls for retrospective grade bias (the cognitive tendency to rate an exam negatively primarily due to receiving an unsatisfactory grade).
Theoretical Framework
The theoretical architecture of the MFE-K integrates foundational tenets from higher education didactics, psychometric test theory, cognitive load theory, and the psychology of academic achievement:
Higher Education Didactics and Quality Assurance
Following the seminal work of Heiner Rindermann (1996) on teaching evaluation in higher education, institutional evaluations serve multiple distinct functions: improving individual teaching competencies, uncovering systemic curricular weaknesses, structuring dialogue between students and faculty, informing resource allocation, and guiding educational development programs. While early university evaluations focused almost exclusively on instructional communication, Dany, Szczyrba, and Wildt (2008) highlighted that testing procedures represent an inseparable, defining component of academic instruction. Exams dictate the "hidden curriculum"; according to the testing effect and assessment-driven learning frameworks, the design, cognitive depth, and perceived fairness of examinations exert a more profound impact on how students study than the passive delivery of lecture material.
Constructive Alignment and Criterion-Referenced Measurement
The MFE-K is heavily anchored in John Biggs’ theory of constructive alignment (Biggs, 1996). Under this paradigm, effective educational systems maintain rigorous concordance between intended learning outcomes (ILOs), the teaching and learning activities (TLAs) conducted in lectures or seminars, and the assessment tasks used to gauge mastery. When exams introduce unexpected cognitive operations (e.g., asking for trivia memorization when critical thinking was taught, or demanding complex problem solving when only superficial definitions were presented), constructive alignment breaks down. The MFE-K empirically verifies whether students perceive this alignment to be intact.
Qualitative Conceptions of "Good Examinations"
Because higher education literature historically lacked formalized theoretical criteria for what constitutes a high-quality written university exam, Froncek and Thielsch (2011; see also Froncek, 2010) executed qualitative interview studies using Mayring’s qualitative content analysis. Conducting semi-structured interviews with both university lecturers across academic disciplines and psychology students across various semester tiers, they inductively distilled key categories of exam excellence: transparency of expectations, precision of instructional prompts, authentic problem-solving orientation, absence of arbitrary catch questions, appropriate pacing, and supportive testing environments. These empirically derived categories directly informed the operationalization and factor design of the MFE-K.
Validity
The validity of the MFE-K has been established across multiple validation cohorts at the Westfälische Wilhelms-Universität Münster and affiliated institutions, examining content, construct, and criterion-related validity.
Content Validity
In educational measurement, establishing content validity for course and exam evaluations is notoriously challenging due to the historical absence of universal external quality benchmarks (Marsh, 1984). The MFE-K established content validity through rigorous multi-stage inductive-deductive construction. Initial items were derived from student council archive records, institutional course evaluation batteries, and Swiss federal technical testing guidelines (Eugster & Lutz, 2004). This initial pool underwent extensive expert reviews involving university examiners, academic commissions, and student representatives. Subsequent qualitative interview studies (Froncek & Thielsch, 2011) systematically categorized instructor and student mental models of exam quality, ensuring that every retained item directly mapped onto verified educational quality dimensions.
Construct and Factorial Validity
Construct validity is evidenced through iterative factor-analytic testing across successive cohorts. Early two-factor formulations demonstrated suboptimal confirmatory fit (χ2 = 146.46, df = 26; TLI = .87, CFI = .91, RMSEA = .11). Subsequent refinements led to well-fitting multi-dimensional models. A three-factor confirmatory model evaluated on N = 688 test protocols across 17 distinct exams demonstrated acceptable goodness-of-fit indices: χ2 = 107.83, df = 41; TLI = .96, CFI = .97, RMSEA = .06. Inter-scale latent correlations confirmed meaningful construct distinctions: Student Burden correlated negatively with Exam Design (r = -.27) and Transparency (r = -.30), whereas Exam Design and Transparency exhibited a robust positive association (r = .62).
Criterion-Related and Discriminant Validity
To establish that the instrument measures true variation between academic assessments rather than homogenous response sets or invariant institutional dissatisfaction, multivariate analyses of variance (MANOVA) were conducted using individual examination modules as the independent factor and MFE-K subscale scores as dependent criteria. The analysis demonstrated significant, robust differentiation across examinations (F = 11.9, df = 48, p < .01, η2 = .22). This substantial effect size indicates that 22% of the total variance in student evaluation scores was attributable to actual differences between specific exams, confirming the instrument’s high diagnostic sensitivity to divergent testing conditions, exam formats, and instructional preparation.
Reliability
The reliability of the MFE-K scales has been evaluated across extensive cross-sectional implementations, demonstrating solid to high internal consistency despite the deliberately parsimonious length of individual subscales:
- Overall Reliability Range: Across primary validation studies involving N = 688 respondents evaluating 17 examinations, subscale Cronbach’s alpha coefficients ranged from .76 to .84.
- Student Burden Subscale (3 items): Cronbach’s alpha reached α = .79 (mean item-total correlations ranged from rit = .48 to .73).
- Transparency Subscale (4 items): Cronbach’s alpha reached α = .84 (item-total correlations ranged from rit = .59 to .76).
- Exam Construction and Design Subscale (4 items): Cronbach’s alpha reached α = .76 (item-total correlations ranged from rit = .47 to .65).
- Extended Preparation and Framework Dimensions: Multi-item modules assessing preparation depth and environmental adequacy similarly demonstrate acceptable internal consistencies typically exceeding α = .75.
Because the questionnaire is intentionally administered immediately following examination completion prior to grade disclosure, test-retest reliability across long intervals is theoretically inappropriate: examinee retrospective evaluations shift fundamentally once formal grades are received. However, split-half reliability and split-sample stability across parallel examination cohorts have repeatedly shown stable metric properties across academic semesters.
Factor Analysis
The factorial structure of the MFE-K was established through a series of exploratory (EFA) and confirmatory factor analyses (CFA) conducted across four consecutive academic semesters (Winter Semester 2007/2008 through Winter Semester 2010/2011). Maximum Likelihood Estimation within structural equation modeling software (IBM SPSS AMOS) served as the standard estimation technique.
Structural Revisions and Model Progression
Early drafts in 2008 utilized an exploratory two-factor solution comprising "Exam and Organization" and "Student Burden." When tested confirmatorily in 2009 (N = 409), this model proved inadequate due to elevated residual covariances (RMSEA = .11, TLI = .87). The subsequent integration of qualitative interview findings (Froncek & Thielsch, 2011) led to a distinct separation of transparency, organizational framework, and specific question design.
Confirmatory Factor Model Parameters
The established model structure demonstrates excellent psychometric fit indices across benchmark criteria (N = 688):
- Model Fit Statistics: χ2 = 107.83, degrees of freedom (df) = 41; Comparative Fit Index (CFI) = .97; Tucker-Lewis Index (TLI) = .96; Root Mean Square Error of Approximation (RMSEA) = .06.
- Standardized Factor Loadings (λ):
- Belastung der Studierenden (Student Burden): Item 1 (λ = .84), Item 2 (λ = .91), Item 3 (λ = .51).
- Transparenz (Transparency): Item 7 (λ = .92), Item 8 (λ = .77), Item 9 (λ = .65), Item 10 (λ = .62).
- Klausurgestaltung (Exam Design): Item 12 (λ = .55), Item 13 (λ = .89), Item 14 (λ = .69), Item 15 (λ = .46).
These robust factor loadings substantiate the structural integrity of the instrument, validating its use as both a set of distinct diagnostic subscale scores and an integrated evaluation battery.
Instrument / Measurement Tool
The MFE-K is a standardized, self-report paper-and-pencil or digital survey instrument designed for immediate administration upon completion of a written university exam.
Structural Parameters
- Test Type: Multi-dimensional student evaluation and academic quality diagnostic tool.
- Target Population: Undergraduate and graduate university students across all academic disciplines undertaking written examinations.
- Total Number of Items: 27 items, categorized into distinct evaluative subscales, diagnostic indicators, and qualitative prompts.
- Administration Time: Approximately 5 to 10 minutes.
- Administration Protocol: Handed out by examination proctors immediately upon submission of the exam paper. Students are instructed to complete the questionnaire prior to leaving the examination hall or return it to a designated secure drop box within a specified window (e.g., up to 14 days, though immediate on-site completion is strongly recommended to preserve affective and cognitive recall). Under no circumstances are student names or matriculation numbers recorded, guaranteeing complete anonymity.
Response Format
The instrument primarily employs a 7-point Likert-type agreement scale formulated in standard German:
- 1 = "stimme gar nicht zu" (strongly disagree)
- 2 = "stimme nicht zu" (disagree)
- 3 = "stimme eher nicht zu" (somewhat disagree)
- 4 = "neutral" (neutral)
- 5 = "stimme eher zu" (somewhat agree)
- 6 = "stimme zu" (agree)
- 7 = "stimme vollkommen zu" (strongly agree)
Specific supplementary items utilize alternative response modes: open numerical entry (e.g., hours spent studying, expected grade), German academic grading scale (1 = sehr gut / very good to 6 = ungenügend / insufficient), or open-ended text fields for constructive feedback.
Scoring and Subscale Architecture
Scale scores are computed by summing or averaging the individual item responses within each subscale. Reverse-scoring or inverted interpretation rules apply to negative indicators:
- Skala 1: Angemessenheit und Güte der Klausur (Exam Adequacy and Construction Quality): Encompasses Items 1 through 9. Higher scores denote superior didactic construction, question clarity, fair difficulty calibration, and authentic comprehension testing.
- Skala 2: Übereinstimmung mit der Vorlesung (Curricular Alignment with Lecture): Encompasses Items 10 through 13. Higher scores indicate strong constructive alignment between taught and examined content.
- Skala 3: Rahmenbedingungen (Environmental and Administrative Framework): Encompasses Items 14 through 18. Higher scores reflect quiet, professional, comfortable, and technically sound examination settings.
- Skala 4: Vorbereitung (Student Preparation): Encompasses Items 19 through 23. Higher scores capture thorough, well-structured student revision and high invested effort.
- Skala 5: Notenerwartung und Gesamteindruck (Grade Expectation and Overall Appraisal): Encompasses Items 24 through 27. Includes quantitative hours studied (Item 24), anticipated grade (Item 25), global academic school mark from 1 to 6 (Item 26; where 1 represents the highest possible quality and 6 represents failing quality), and qualitative open feedback (Item 27).
- Subscale Interpretation Note: On scales assessing student burden or difficulty, lower to intermediate scores are typically targeted for optimal curriculum balance, whereas high scores on transparency and exam design reflect high instructional quality.
Permissions & Fee and Test Year
The Münster Questionnaire for the Evaluation of Exams was initially conceptualized in 2007 and formally validated through subsequent iterative cohorts culminating in standard operational deployment in 2011. The instrument was developed within the Department of Psychology at the Westfälische Wilhelms-Universität Münster as an open-access public-science contribution to university quality assurance. The scale is non-commercial and can be utilized free of charge by academic institutions, university departments, and researchers for educational and scientific evaluation purposes, provided proper scholarly citation is maintained. Adaptation or integration into digital campus evaluation platforms (e.g., EvaSys, LimeSurvey) is permitted. Inquiries regarding original institutional benchmarks or institutional documentation may be directed to Dr. Meinald T. Thielsch or Benjamin Froncek.
References
- Bechler, A., & Thielsch, M. T. (2012). Schwierigkeiten von Studierenden in der Prüfungsvorbereitung [Students’ difficulties in examination preparation]. Vortrag auf dem 48. Kongress der Deutschen Gesellschaft für Psychologie (DGPs), Bielefeld.
- Biggs, J. (1996). Enhancing teaching through constructive alignment. Higher Education, 32(3), 347–364. https://doi.org/10.1007/BF00138871
- Dany, S., Szczyrba, B., & Wildt, J. (Eds.). (2008). Prüfungen auf die Agenda! Hochschuldidaktische Perspektiven auf Reformen im Prüfungswesen [Exams on the agenda! Higher education didactic perspectives on testing reforms]. W. Bertelsmann Verlag.
- Eugster, B., & Lutz, B. (2004). Leitfaden für das Planen, Durchführen und Auswerten von Prüfungen an der ETHZ [Guidelines for planning, executing, and evaluating examinations at ETH Zurich]. Didaktikzentrum der ETH Zürich.
- Froncek, B. (2010). Entwicklung und Validierung eines Fragebogens zur Evaluation universitärer Prüfungen [Development and validation of a questionnaire for the evaluation of university examinations] (Unpublished master’s thesis). Westfälische Wilhelms-Universität Münster.
- Froncek, B., & Thielsch, M. T. (2011). Merkmale guter schriftlicher Prüfungen aus studentischer und Prüfersicht [Characteristics of good written examinations from student and examiner perspectives]. Zeitschrift für Hochschulentwicklung, 6(3), 193–208. https://doi.org/10.3217/zfhe-6-03/13
- Kultusministerkonferenz. (2005). Handreichung für die Akkreditierung von Bachelor- und Masterstudiengängen [Handout for the accreditation of bachelor’s and master’s degree programs]. KMK.
- Marsh, H. W. (1984). Students’ evaluations of university teaching: Dimensionality, reliability, validity, potential biases, and utility. Journal of Educational Psychology, 76(5), 707–754. https://doi.org/10.1037/0022-0663.76.5.707
- Mayring, P. (2003). Qualitative Inhaltsanalyse: Grundlagen und Techniken [Qualitative content analysis: Fundamentals and techniques] (8th rev. ed.). Beltz.
- Rindermann, H. (1996). Untersuchungen zur Validität von Studentischen Lehrveranstaltungsevaluationen [Studies on the validity of student course evaluations]. Beltz.
- Thielsch, M. T., Froncek, B., & Hertel, G. (2008). Evaluation von Prüfungen im Rahmen von Bachelor-Studiengängen [Evaluation of examinations in bachelor’s degree programs]. Vortrag auf dem 46. Kongress der Deutschen Gesellschaft für Psychologie (DGPs), Jena.
- Thielsch, M. T., Froncek, B., & Hertel, G. (2010). Prüfungsevaluation im Bachelor Psychologie: Drei Semester Praxistest [Exam evaluation in the psychology bachelor’s program: Three semesters of field testing]. Poster präsentiert auf dem 47. Kongress der Deutschen Gesellschaft für Psychologie (DGPs), Bremen.
Items of the Scale
Antwortskala (Items 1–23):
Verwendet wird weitgehend ein 7-stufiges Antwortformat mit den Optionen:
1 = "stimme gar nicht zu"
2 = "stimme nicht zu"
3 = "stimme eher nicht zu"
4 = "neutral"
5 = "stimme eher zu"
6 = "stimme zu"
7 = "stimme vollkommen zu"
Skala 1: Angemessenheit und Güte der Klausur
- Die Klausurfragen waren verständlich formuliert.
- Die Klausur war angemessen schwierig.
- Die Klausur war fair gestaltet.
- Die Klausur erforderte ein echtes Verständnis des Stoffes (nicht bloßes Auswendiglernen).
- Bei den Aufgaben war klar, wie viel sie zur Gesamtnote beitragen.
- Die Bearbeitungszeit für die Klausur war ausreichend bemessen.
- Die Punkteverteilung auf die einzelnen Aufgaben war angemessen.
- Die Klausur stellte eine repräsentative Stichprobe des Stoffes dar.
- Insgesamt beurteile ich die Klausur als gut konstruiert.
Skala 2: Übereinstimmung mit der Vorlesung
- Die Inhalte der Klausur stimmten gut mit den Inhalten der Lehrveranstaltung überein.
- Der in der Vorlesung gesetzte Schwerpunkt spiegelte sich in der Klausur wider.
- Die Aufgaben bezogen sich auf den Stoff, der auch im Unterricht behandelt wurde.
- Es gab keine Überraschungen hinsichtlich der abgefragten Themenbereiche.
Skala 3: Rahmenbedingungen
- Die räumlichen Bedingungen (z.B. Akustik, Beleuchtung, Platzangebot) waren angemessen.
- Die Prüfungsaufsicht sorgte für eine ruhige und ungestörte Atmosphäre.
- Die organisatorischen Abläufe vor und während der Klausur funktionierten reibungslos.
- Fragen zur Aufgabenstellung während der Klausur wurden angemessen beantwortet.
- Die technischen Voraussetzungen (z.B. Lesbarkeit des Drucks, ggf. PC-Systeme) waren einwandfrei.
Skala 4: Vorbereitung
- Ich habe mich intensiv auf diese Klausur vorbereitet.
- Der zeitliche Aufwand für meine Prüfungsvorbereitung war hoch.
- Ich habe den relevanten Vorlesungsstoff vor der Klausur vollständig durchgearbeitet.
- Meine Vorbereitung war zielgerichtet und strukturiert.
- Ich fühle mich durch meine Vorbereitung gut auf diese Klausur vorbereitet.
Skala 5: Notenerwartung und Gesamteindruck
- Wie viele Stunden haben Sie insgesamt für die Vorbereitung auf diese Klausur aufgewendet? (offene Angabe in Stunden)
- Welche Note erwarten Sie in dieser Klausur? (erwartete Prüfungsnote)
- Welche Gesamtnote würden Sie dieser Klausur geben? (Schulnoten 1 = sehr gut bis 6 = ungenügend)
- Was möchten Sie den Prüfenden noch mitteilen (Lob, Kritik, Anregungen)? (Freitext)