1. Abstract
The Münster Questionnaire for the Evaluation of Seminars (German: Münsteraner Fragebogen zur Evaluation von Seminaren, abbreviated as MFE-S) is an economical, multidimensional psychometric instrument developed at the University of Münster designed to assess the instructional quality of academic seminars in higher education. Derived from an initial 17-item evaluation pool and refined into a modular system of course ratings, the operational core of the MFE-S comprises 9 primary psychometric items distributed equally across three distinct dimensions of instructional effectiveness: Didactics (Didaktik; 3 items), Difficulty (Schwierigkeit; 3 items), and Self-Study (Selbststudium; 3 items). Five additional administrative and evaluative single items capture attendance motives, pre-existing subject interest, subjective knowledge gain, a holistic grade on a standard 0–15 academic point scale, and qualitative open-ended feedback, bringing the full instrument to 14 items.
Items within the three core subscales are rated along a 7-point Likert scale anchored from 1 (completely inaccurate / völlig unzutreffend) to 7 (completely accurate / völlig zutreffend). Psychometric evaluation conducted on a rigorously filtered independent sample of university students ($N = 103$, drawn from a broader repository of 9,757 evaluations collected between the winter semester 2002/2003 and summer semester 2008) confirmed a three-factor oblique structural model via confirmatory factor analysis (CFA). The model displayed excellent fit to the empirical data ($\chi^2 = 31$, $df = 24$, $p = .155$; $ ext{CFI} = .98$;$ ext{TLI} = .97$;$ ext{RMSEA} = .05$). Internal consistency reliability estimates (Cronbach’s alpha) reflect deliberate bandwidth-fidelity considerations across the ultra-short subscales: $lpha = .86$ for Didactics, $lpha = .62$ for Difficulty, and $lpha = .57$ for Self-Study. The MFE-S addresses critical methodological limitations inherent in conventional, lengthy student evaluation of teaching (SET) batteries by drastically reducing respondent burden, preserving high response rates ($>60%$ in non-mandatory settings), and delivering actionable diagnostic profiles for tertiary-level educators and university quality assurance units.
2. Keywords
student evaluation of teaching, higher education assessment, course evaluation, instructional quality, seminar pedagogy, psychometrics, confirmatory factor analysis, didactics, self-regulated learning, academic quality assurance, Münster questionnaire for the evaluation of seminars, MFE-S
3. Authors
The Münster Questionnaire for the Evaluation of Seminars was developed, validated, and implemented by researchers at the Department of Psychology at the University of Münster (Westfälische Wilhelms-Universität Münster), Germany:
- Dr. Dipl.-Psych. Meinald T. Thielsch — University of Münster, Institute of Psychology (Psychologisches Institut 1), Fliednerstraße 21, 48149 Münster, Germany. Email: [email protected]. Primary research areas: Human-Computer Interaction, online assessment, student evaluations of teaching, and organizational psychology.
- Dr. Dipl.-Psych. Gerrit Hirschfeld — Vodafone Stiftungsinstitut und Lehrstuhl für Kinderschmerztherapie und Pädiatrische Palliativmedizin, Vestische Kinder- und Jugendklinik Datteln, Witten/Herdecke University, Dr.-Friedrich-Steiner-Straße 5, 45711 Datteln, Germany. Email: [email protected]. Primary research areas: Quantitative methodology, advanced psychometrics, pediatric pain, and educational measurement.
- Collaborating Contributors & Foundation Authors: Additional scale conceptualization and primary item pool generation were contributed by Dipl.-Psych. Jens Grabbe (original 17-item scale development, 2003), Dipl.-Psych. Ingo Haaser (item reduction and psychometric calibration, 2006), and Dipl.-Psych. Claudia Moeck (implementation of the automated web-based feedback system, 2007).
4. Purpose
The primary purpose of the Münster Questionnaire for the Evaluation of Seminars (MFE-S) is to provide an empirically grounded, economically streamlined, and psychometrically robust diagnostic instrument for evaluating teaching effectiveness in university seminars. Student evaluations of educational quality represent an essential cornerstone of academic quality management across contemporary European and international tertiary institutions. Grounded in the theoretical framework established by Rindermann (1996, 2001), institutional teaching evaluations serve multiple concurrent functions: facilitating instructional self-reflection and professional pedagogical growth among faculty, detecting curricular strengths and pedagogical deficits at course, departmental, and institutional levels, stimulating productive academic dialogue between instructors and students, informing institutional allocation of instructional resources, and systematically monitoring faculty development interventions.
Prior to the introduction of the MFE-S, prevailing German-language instruments designed to capture student ratings of instruction—such as the Fragebogen zur Evaluation von Lehrveranstaltungen (FEVOR; Staufenbiel, 2000), the Heidelberger Inventar zur Lehrveranstaltungsevaluation (HILVE; Rindermann, 2001), the Kieler Evaluationsinventar für Lehrveranstaltungen (KIEL; Gediga et al., 2000), and the Trierer Inventar zur Lehrevaluation (TRIL; Gollwitzer & Schlotz, 2003)—featured comprehensive item banks extending across 30 to over 60 items. Although methodologically thorough, the sheer length of these traditional instruments presents severe structural impediments when deployed within modern digital campus architectures. When students are asked to evaluate five to eight separate courses concurrently at the conclusion of an academic semester, lengthy inventories trigger respondent fatigue, rapid non-differentiated response sets (straightlining), elevated attrition rates, and severe non-response bias. Consequently, response rates in un-incentivized, voluntary online evaluations historically collapsed below representative thresholds, often yielding feedback that reflected only the extreme positive or negative margins of the student cohort.
The MFE-S was systematically engineered to resolve this trade-off between psychometric integrity and survey economy. Specifically, the instrument was crafted for integration into an automated, web-based quality management architecture capable of simultaneous administration across diverse seminar offerings. In higher education pedagogy, a seminar differs fundamentally from a formal didactic lecture. Whereas lectures emphasize expository instruction, unilateral information transfer, and broad conceptual surveys delivered to large audiences, seminars demand interactive discourse, collaborative inquiry, dynamic problem-solving, cognitive challenge, and extensive individual preparation (self-study). The MFE-S intentionally eschews lecture-centric metrics to isolate three specific instructional levers governing seminar success: the instructor’s didactical and communicative competence, the calibration of intellectual difficulty to student baseline knowledge, and the adequacy of pedagogical scaffolding supporting student self-regulated preparation.
Beyond individual formative feedback, the MFE-S serves rigorous institutional research purposes. Its web-based infrastructure allows automated norm-referenced comparisons, enabling instructors to benchmark their seminar scores against historical department-wide baselines within equivalent course categories. In institutional accreditation and pedagogical scholarship, the standardized structure of the MFE-S allows longitudinal tracking of curricular reforms and provides clear diagnostic indicators identifying whether suboptimal course ratings stem from communicative friction (didactics), cognitive mismatch (pacing/difficulty), or insufficient instructional resources (self-study materials).
5. Psychological Construct
The overarching target construct of the MFE-S is perceived seminar quality (multidimensional instructional effectiveness in interactive higher education contexts). In accordance with modern consensus within educational psychology, instructional effectiveness is not a monolithic, unitary entity, but a hierarchical, multidimensional construct comprising distinct interactive behaviors, instructional designs, and cognitive demands (Centra, 1993; Feldman, 1989; Marsh, 1984, 2007). The MFE-S operationalizes seminar quality through three correlated primary latent dimensions, complemented by selective single-item contextual indicators:
1. Didactics (Didaktik)
The Didactics dimension reflects the instructor’s communicative clarity, pedagogical methodology, perceived teaching engagement, and instructional enthusiasm. Within an interactive seminar environment, didactic excellence extends beyond clear speech; it requires the cognitive translation of abstract theoretical concepts into accessible, intellectually digestible components while fostering an engaging scholarly atmosphere. In the MFE-S, this construct is tapped by three specific operational manifestations:
- Explanatory Clarity: The degree to which the instructor successfully elucidates complex, challenging subject matter in a manner that is lucid and comprehensible (Item 3: “Der/Die Lehrende erläuterte schwierige Sachverhalte verständlich”).
- Instructional Commitment: The perceptible intrinsic value and priority the instructor places on the pedagogical process, communicating dedication and professional investment in student learning (Item 4: “Man hat gemerkt, dass der Dozent/die Dozentin die Lehre für wichtig hält”).
- Methodological Appropriateness: The perceived efficacy, fitness-for-purpose, and didactic alignment of the selected teaching methods relative to the curriculum being taught (Item 5: “Die Lehrmethoden waren zur Vermittlung des Stoffes gut geeignet”).
High scores on Didactics denote an instructor who commands high explanatory precision, projects professional engagement, and utilizes teaching strategies that directly facilitate conceptual comprehension.
2. Difficulty (Schwierigkeit)
The Difficulty dimension captures the intellectual calibration, cognitive pacing, and methodological variance of the seminar relative to the student cohort’s baseline competencies. Unlike lecture formats where difficulty is largely determined by content density, seminar difficulty is dynamically co-constructed through classroom discussions, reading volume, and assigned analytical tasks. Grounded in cognitive load theory and Vygotsky’s zone of proximal development, effective instruction must strike a balance between cognitive underload (inducing boredom) and excessive cognitive load (inducing frustration and disengagement). The MFE-S models this construct via three items:
- Cognitive Calibration: The instructor’s active adjustment of the seminar’s intellectual depth and pacing to harmonize with the empirical baseline knowledge of the enrolled students (Item 6: “Der/Die Lehrende passte das Niveau des Seminars an den Wissensstand der Studierenden an”).
- Subjective Overload: An indicator of cognitive strain, measuring whether the seminar content exceeded the student’s individual processing capacity or prerequisite knowledge (Item 7: “Die Inhalte des Seminars waren zu schwierig für mich”).
- Methodological Monotony: An indicator reflecting pedagogical variety versus experiential fatigue, assessing whether instructional techniques lacked diversification (Item 8: “Die Lehrmethoden des Seminars waren wenig abwechslungsreich”).
3. Self-Study (Selbststudium)
The Self-Study dimension assesses the structural and material scaffolding provided by the academic department or instructor to foster autonomous learning, combined with the student’s personal investment in course preparation. In higher education seminars, substantive cognitive gains occur predominantly outside the physical classroom through independent literature analysis, preparation of discussion briefs, and problem sets. The MFE-S conceptualizes this dimension as an interaction between pedagogical resources and student self-regulation:
- Quantitative Resource Adequacy: The sufficient availability of supplementary educational materials, literature recommendations, and reading packets (Item 9: “Es wurden ausreichend Materialien (z.B. Literaturangaben, Skript) zur Vertiefung des Stoffes angeboten”).
- Qualitative Resource Utility: The pragmatic, pedagogical helpfulness and clarity of the provided literature and study guides in deepening conceptual mastery (Item 10: “Es wurden hilfreiche Materialien (z.B. Literaturangaben, Skript) zur Vertiefung des Stoffes angeboten”).
- Student Behavioral Investment: The self-reported frequency and regularity with which the student actively prepared for seminar sessions via reading assigned texts or completing coursework (Item 11: “Ich habe die Seminarsitzungen regelmäßig vorbereitet (z.B. durch das Lesen von Literatur oder die Bearbeitung von Hausaufgaben)”).
Supplementary Diagnostic Indicators
In addition to the nine latent core items, the MFE-S incorporates five standalone contextual items that do not load on the structural factors but provide vital control parameters and holistic criteria:
- Attendance Motivation (Item 1): Categorical assessment classifying whether enrollment was driven by mandatory curricular requirements (Pflicht), intrinsic academic interest (Interesse), or external circumstantial factors (Sonstiges).
- Prior Subject Interest (Item 2): A continuous 7-point Likert rating measuring pre-existing fascination with the seminar topic prior to course commencement, essential for controlling student selection bias.
- Subjective Competence Gain (Item 12): A 7-point Likert item assessing perceived learning success and content acquisition (“Ich habe in der Veranstaltung inhaltlich viel gelernt”).
- Global Performance Mark (Item 13): A holistic evaluation scored on the standardized German upper-secondary school grading metric (gymnasiale Oberstufe; 0 to 15 points, where 0 represents ungenügend [failing] and 15 represents sehr gut + [flawless distinction]).
- Formative Qualitative Feedback (Item 14): An open-ended text box capturing unstructured student remarks regarding highlights, targeted suggestions for structural reform, and personal feedback.
6. Theoretical Framework
The conceptual architecture of the MFE-S is grounded in the Multimodal Model of Teaching Quality and Success (Multimodales Bedingungsmodell des Lehrerfolgs) formulated by educational psychometrician Heiner Rindermann (2001), alongside the seminal multidimensional frameworks of Herbert W. Marsh (1984, 2007) and John A. Centra (1993). Rindermann’s model posits that teaching success in tertiary education is an emergent product of a complex, reciprocal system comprising four primary determinant clusters:
- Instructor Attributes: Encompassing subject matter expertise, pedagogical content knowledge, communication skills, perceived enthusiasm, and structural didactic organization.
- Student Preconditions: Encompassing baseline cognitive ability, prior subject-matter knowledge, intrinsic motivation, and academic self-efficacy.
- Contextual / Environmental Conditions: Class size, departmental infrastructure, physical learning environments, technical learning systems, and curricular integration.
- Interactive Instructional Process: The dynamic interplay of cognitive demand, mutual communication, and independent self-study occurring over the duration of the course.
Classical educational measurement literature frequently debates whether student ratings reflect objective teaching effectiveness or merely subjective affective popularity (the “Dr. Fox effect”). Marsh’s extensive empirical work (1984, 1987) demonstrated that student evaluations of educational quality are inherently multidimensional; treating course evaluation as a unifactorial “good vs. bad” index collapses crucial diagnostic distinctions and inflates halo error. An instructor may possess exceptional didactic clarity while assigning an excessively demanding cognitive workload, or conversely, an instructor may establish an effortless, entertaining seminar dynamic that yields negligible objective learning gains. By parsing evaluation into distinct operational vectors (Didactics, Difficulty, and Self-Study), the MFE-S aligns with this multidimensional imperative.
Furthermore, the design of the MFE-S explicitly integrates principles from Cognitive Load Theory (Sweller, 1988; Paas et al., 2003). In university seminars, learning materials present substantial intrinsic cognitive load. If the instructor’s didactic delivery is disorganized or conceptually obscure, students experience elevated extraneous cognitive load, which impairs working memory resources necessary for deep conceptual integration (germane cognitive load). The MFE-S’s Didactics scale directly indexes the instructor’s ability to minimize extraneous load through lucid explanations and fit-for-purpose methods. Concurrently, the Difficulty scale monitors the alignment between intrinsic content complexity and student prior knowledge, verifying that the instruction resides within the optimal zone of proximal development without inducing cognitive overload.
Finally, the scale incorporates modern paradigms of self-regulated learning (Zimmerman, 2000; Boekaerts, 1999). Academic seminars are not passive instructional experiences; knowledge construction depends directly upon students’ autonomous self-study behaviors. In the MFE-S, the Self-Study scale models both the external institutional scaffolding (materials, literature packets, reading syllabi) and the student’s personal behavioral compliance (active reading and homework preparation), acknowledging that student agency is an inseparable determinant of seminar success.
7. Validity
Validity evidence for the MFE-S derives from content-oriented, construct-related, and criterion-referenced empirical investigations conducted across multi-year implementation phases within the Department of Psychology at the University of Münster.
Content Validity
Content validity was established through systematic iterative pruning of an extensive 17-item pedagogical inventory developed by Grabbe (2003). In alignment with Rindermann’s (2001) empirical taxonomy of instructional success, the two most decisive instructor-controlled predictors of academic outcomes are didactic communication and calibrated task difficulty. The MFE-S developers specifically isolated items tapping these core functional mechanisms while intentionally excluding non-instructional trivialities (e.g., room temperature, technical projector malfunctions, personality quirks). The three items composing the Didactics subscale represent the quintessential triad of pedagogical clarity, professional commitment, and methodological suitability, ensuring robust domain coverage of instructional competence.
Construct and Factorial Validity
Construct validity is substantiated by confirmatory factor analytic testing of the hypothesized three-factor structural model. In a validation sample of $N = 103$ independent student raters (rigorously filtered to eliminate duplicate entries and preserve the mathematical assumption of observation independence), the three-factor oblique measurement model demonstrated superior fit ($\chi^2 = 31$, $df = 24$, $p = .155$; $ ext{TLI} = .97$;$ ext{CFI} = .98$;$ ext{RMSEA} = .05$). The empirical factor loadings confirm that the selected items serve as valid structural indicators of their designated latent targets:
- Didactics item loadings range from $lambda = .72$ to $lambda = .87$, signifying strong convergent validity on the underlying pedagogical communication factor.
- Difficulty item loadings display notable divergence: Item 6 (calibrating level to knowledge) loaded at $lambda = .84$ and Item 7 (excessive content difficulty) loaded at $lambda = .71$, demonstrating strong representation of cognitive challenge. Item 8 (lack of methodological variety) loaded moderately at $lambda = .33$, reflecting its dual association with pacing and instructional engagement.
- Self-Study item loadings demonstrate high convergent loading for institutional material provision (Item 9: $lambda = .74$; Item 10: $lambda = .96$), whereas individual preparation behavior (Item 11) exhibited a lower structural loading ($lambda = .18$), capturing student behavioral variance that operates semi-independently of material provision.
Discriminant and Inter-Subscale Validity
Inter-scale correlation analyses (Table 2 in source documentation) provide clear evidence of discriminant validity between the three latent dimensions:
- Didactics and Difficulty correlate positively and moderately at $r = .66$ ($p < .001$), reflecting that superior didactical clarity is associated with perceived appropriateness of course difficulty.
- Didactics and Self-Study share a modest correlation of $r = .34$, indicating that an instructor’s explanatory skill operates distinctly from whether students receive adequate reading packets or invest time in independent preparation.
- Difficulty and Self-Study demonstrate a low correlation of $r = .24$, proving that cognitive difficulty is not a simple linear artifact of homework volume.
Criterion and Social Validity
Social validity and practical utility are corroborated by multi-year institutional meta-evaluations (Haaser et al., 2007; Thielsch et al., 2008). In voluntary, un-incentivized deployment across consecutive semesters, the MFE-S maintained departmental seminar participation rates averaging above $60%$. These figures markedly exceed typical institutional online survey returns, demonstrating high acceptance among both students and faculty. Faculty reported that the three-factor profile provided readily interpretable diagnostic insights that directly guided targeted curriculum revisions.
8. Reliability
The internal consistency of the MFE-S subscales was evaluated via Cronbach’s alpha ($lpha$) utilizing the psychometrically clean calibration sample ($N = 103$). Because each subscale comprises exactly three items designed to maximize domain breadth rather than semantic redundancy, reliability coefficients must be interpreted through the lens of the classical bandwidth-fidelity dilemma (Cronbach, 1960; Streiner, 2003):
- Didactics Subscale ($lpha = .86$): Displays high internal consistency. The item-total correlations ($r_{it}$) are strong across all three indicators: Item 3 ($r_{it} = .70$), Item 4 ($r_{it} = .66$), and Item 5 ($r_{it} = .68$). Removing any single item would decrease the overall scale reliability (alpha-if-item-deleted drops to .82 for Item 3, .85 for Item 4, and .74 for Item 5), verifying that all three items contribute decisively to measuring didactic competence.
- Difficulty Subscale ($lpha = .62$): Demonstrates acceptable internal consistency for an ultra-short 3-item exploratory and formative evaluation tool. Item-total correlations for Item 6 ($r_{it} = .56$) and Item 7 ($r_{it} = .52$) are robust. Item 8 exhibits a lower item-total correlation ($r_{it} = .24$). Deleting Item 8 would increase the subscale alpha to $.76$. However, the scale authors deliberately retained Item 8 to ensure the subscale captures experiential monotony alongside cognitive overload.
- Self-Study Subscale ($lpha = .57$): Yields a moderate internal consistency coefficient. Analysis of individual item metrics reveals that Item 9 (material availability; $r_{it} = .50$) and Item 10 (material helpfulness; $r_{it} = .57$) are strongly intercorrelated ($r = .71$). Deleting Item 11 (student preparation behavior; $r_{it} = .15$) would substantially elevate the alpha coefficient of the remaining two-item dyad to $.83$. The authors retained Item 11 on theoretical grounds: omitting student preparation would transform the subscale into a pure measure of reading list distribution, discarding the interactive reality that self-study requires both resource provision and active student engagement.
These reliability statistics are fully commensurate with benchmark values reported across established full-length German teaching inventories, including the FEVOR (Staufenbiel, 2000), HILVE (Rindermann, 2001), and TRIL (Gollwitzer & Schlotz, 2003), while achieving these metrics at a fraction of the respondent time cost.
9. Factor Analysis
The dimensional structure of the MFE-S was initially explored through exploratory principal component analysis (PCA) on an initial pool of 17 items (Grabbe, 2003; Haaser, 2006). This reduction phase eliminated 8 items that displayed unacceptable cross-loadings or inadequate discriminatory power, yielding a 9-item structure loading clearly across three components representing Didactics, Difficulty, and Self-Study.
To rigorously confirm this structural model without capitalization on chance, a linear Confirmatory Factor Analysis (CFA) was conducted using AMOS. Maximum Likelihood (ML) estimation was employed. A paramount methodological challenge in educational course evaluations involves non-independent observations: individual students frequently evaluate multiple courses, and courses cluster within instructors, violating the assumption of independent and identically distributed observations. To solve this, the researchers created an uncorrelated validation sample by extracting evaluations from the semester with the highest return rate (Winter Semester 2007/2008) and systematically purging all records sharing identical demographic vectors (age, gender, semester, degree track). This yielded a guaranteed independent sample of $N = 103$ distinct student evaluations whose demographic profile perfectly matched the parent population.
Model Fit and Parameter Estimates
The three-factor oblique measurement model specifying simple structure (each item loading exclusively onto its designated latent construct, with latent factors permitted to covary) achieved exceptional goodness-of-fit indices:
- $\chi^2 = 31.00$ ($df = 24$, $p = .155$)
- $\chi^2 / df = 1.29$ (well below the conservative threshold of $2.0$)
- $ ext{CFI (Comparative Fit Index)} = .98$
- $ ext{TLI (Tucker-Lewis Index)} = .97$
- $ ext{RMSEA (Root Mean Square Error of Approximation)} = .050$ ($90% \text{ CI } [.000, .095]$)
The standardized factor loadings ($lambda$), item means ($M$), standard deviations ($SD$), item-total correlations ($r_{it}$), and alpha-if-deleted values ($CA$) are detailed in the psychometric matrix below:
| Subscale & Item Code | Item Text Abbreviated | $M$ | $SD$ | $r_{it}$ | $lambda$ (FL) | $lpha$ if del. |
|---|---|---|---|---|---|---|
| Didactics (D1 / Item 3) | Explanations were understandable | 5.2 | 1.9 | .70 | .87 | .82 |
| Didactics (D2 / Item 4) | Instructor values teaching | 5.6 | 1.8 | .66 | .72 | .85 |
| Didactics (D3 / Item 5) | Methods suited for content | 5.0 | 1.7 | .68 | .87 | .74 |
| Difficulty (SC1 / Item 6) | Pace adapted to baseline knowledge | 4.8 | 2.0 | .56 | .84 | .34 |
| Difficulty (SC2 / Item 7) | Content too difficult for me | 4.4 | 2.1 | .52 | .71 | .38 |
| Difficulty (SC3 / Item 8) | Methods lacked variety | 4.7 | 2.0 | .24 | .33 | .76 |
| Self-Study (SE1 / Item 9) | Sufficient materials provided | 5.1 | 1.7 | .50 | .74 | .30 |
| Self-Study (SE2 / Item 10) | Helpful materials provided | 4.7 | 1.9 | .57 | .96 | .16 |
| Self-Study (SE3 / Item 11) | Regular preparation completed | 3.8 | 2.1 | .15 | .18 | .83 |
Inter-Item Correlation Matrix
The bivariate correlation patterns across all 9 items ($N = 103$) corroborate the observed factor structure. Intracluster correlations within the Didactics dimension are exceptionally strong ($r_{D1,D2} = .58$, $r_{D1,D3} = .74$, $r_{D2,D3} = .69$). Items 9 and 10 within the Self-Study dimension correlate at $r = .71$. Cross-domain correlations between Didactics and Self-Study preparation behavior (Item 11) hover around zero ($r = .00$ to $.03$), empirically demonstrating the structural divergence between instructor communication and student prep compliance.
10. Instrument / Measurement Tool
- Instrument Name: Münster Questionnaire for the Evaluation of Seminars (MFE-S; Münsteraner Fragebogen zur Evaluation von Seminaren).
- Instrument Type: Standardized multidimensional student rating of instruction (SRI / SET) inventory.
- Target Population: Undergraduate and graduate university students participating in higher education seminars, proseminars, advanced seminars (Hauptseminare), and practical colloquia.
- Administration Format: Self-administered digital assessment implemented via an online survey platform (PHP/MySQL architecture), accessible via desktop, tablet, or mobile browser. It can also be administered via traditional paper-and-pencil questionnaires.
- Total Item Count: 14 items total (comprising 9 core psychometric items across three subscales, plus 5 supplementary administrative, diagnostic, and open-ended items).
- Dimensional Structure (Core Items):
- Didactics (Didaktik): 3 items (Items 3, 4, 5).
- Difficulty (Schwierigkeit): 3 items (Items 6, 7, 8).
- Self-Study (Selbststudium): 3 items (Items 9, 10, 11).
- Response Formats:
- Core Items (3–11) & Item 2 & Item 12: 7-point Likert rating scale ranging from 1 = “völlig unzutreffend” (completely inaccurate / disagree) to 7 = “völlig zutreffend” (completely accurate / agree). All intermediate scale points (2, 3, 4, 5, 6) are unlabelled numerical gradations.
- Item 1 (Attendance Motive): Nominal categorical selection with three choices: Pflicht (compulsory curriculum), Interesse (intrinsic interest), or Sonstiges (other reasons).
- Item 13 (Global Evaluation): Discrete integer rating utilizing the standardized German Gymnasium Oberstufe grading scale (0 to 15 points; 0 = ungenügend, 15 = sehr gut +).
- Item 14 (Qualitative Feedback): Open-ended qualitative text field allowing free-form student commentary on seminar highlights and recommendations for instructional adjustment.
- Scoring and Aggregation Procedures:
- Subscale Scores: Given the demonstrated unidimensionality of each 3-item subscale, individual subscale scores can be computed as either a sum score (ranging from 3 to 21) or an unweighted arithmetic mean score (ranging from 1.0 to 7.0). Mean scoring is preferred for institutional reporting to retain direct interpretability against the original 1–7 response continuum.
- Course-Level Aggregation: Individual student scores within a seminar cohort are averaged to compute course-level didactic, difficulty, and self-study profiles. In the Münster institutional deployment, course aggregated means are automatically plotted against faculty-wide normative distribution percentiles.
- Item Inversion Note: Items 7 and 8 capture negative aspects of course experience (excessive difficulty and lack of methodological variety). Depending on institutional reporting preferences, these items may be reverse-coded (Recode: $8 – X$) if an aggregated composite where higher scores universally designate positive outcomes is desired, or maintained in their raw orientation for specific diagnostic tracking of overload and monotony.
- Completion Time: Approximately 2 to 3 minutes for the core ratings; 4 to 5 minutes when completing open-ended qualitative comments.
11. Permissions & Fee and Test Year
- Test Development History: Developed between 2002 and 2008 at the University of Münster. The initial 17-item version was constructed in 2002/2003 (Grabbe, 2003); psychometric reduction to 14 items occurred in 2005/2006 (Haaser, 2006); automated online deployment was validated in 2007 (Haaser, Thielsch & Moeck, 2007); the definitive confirmatory factor model was formally documented in 2008.
- Publication and Archival Status: The MFE-S is documented in the open-access psychometric repository ZIS – GESIS Leibniz Institute for the Social Sciences (Zusammenstellung sozialwissenschaftlicher Items und Skalen).
- Permissions and Licensing: The instrument is open-access for non-commercial academic research, pedagogical evaluation, and institutional quality assurance in higher education. Researchers and educational practitioners may utilize and adapt the scale without paying licensing fees, provided that appropriate academic attribution and formal citation are given to the scale authors and the University of Münster.
- Commercial Use: Any commercial deployment, integration into proprietary corporate training platforms, or resale within commercial software packages requires prior formal written authorization from the primary authors (Dr. Meinald T. Thielsch).
12. References
- Boekaerts, M. (1999). Self-regulated learning: Where we are today. International Journal of Educational Research, 31(6), 445–457. https://doi.org/10.1016/S0883-0355(99)00014-2
- Centra, J. A. (1993). Reflective faculty evaluation: Enhancing teaching and determining faculty effectiveness. Jossey-Bass.
- Cronbach, L. J. (1960). Essentials of psychological testing (2nd ed.). Harper & Row.
- Feldman, K. A. (1989). The association between student ratings of specific instructional dimensions and student achievement: Refining and extending the synthesis of data from multisection courses. Research in Higher Education, 30(6), 583–645. https://doi.org/10.1007/BF00992392
- Gediga, G., Schömann, K. D., & Hamborg, K. C. (2000). Das Kieler Evaluationsinventar für Lehrveranstaltungen (KIEL). Universität Osnabrück.
- Gollwitzer, M., & Schlotz, W. (2003). Das Trierer Inventar zur Lehrevaluation (TRIL): Ein theoriegeleitetes Instrument zur Messung von Lehrqualität. Diagnostica, 49(1), 34–44. https://doi.org/10.1026//0012-1924.49.1.34
- Grabbe, J. (2003). Entwicklung und Erprobung eines Fragebogens zur Evaluation von Lehrveranstaltungen am Fachbereich Psychologie (Unpublished diploma thesis). Westfälische Wilhelms-Universität Münster, Germany.
- Gritz, W., Soucek, R., & Bacher, M. (2005). Online-Lehrevaluation: Konzeption und Umsetzung eines automatisierten Evaluationssystems. In G. Krampen & H. Zayer (Eds.), Psychologiedidaktik und Evaluation V (pp. 231–242). Hogrefe.
- Haaser, I. (2006). Lehrevaluation an der Universität Münster: Weiterentwicklung und psychometrische Überprüfung der Evaluationsfragebögen (Unpublished diploma thesis). Westfälische Wilhelms-Universität Münster, Germany.
- Haaser, I., Thielsch, M. T., & Moeck, C. (2007). Webbasierte Lehrevaluation in der Praxis: Erfahrungen, Akzeptanz und Rücklaufquoten. In M. Krämer, S. Preiser, & K. Brusdeylins (Eds.), Psychologiedidaktik und Evaluation VI (pp. 317–326). Shaker Verlag.
- Marsh, H. W. (1984). Students’ evaluations of university teaching: Dimensionality, reliability, validity, potential biases, and utility. Journal of Educational Psychology, 76(5), 707–754. https://doi.org/10.1037/0022-0663.76.5.707
- Marsh, H. W. (1987). Students’ evaluations of university teaching: Research findings, methodological issues, and directions for future research. International Journal of Educational Research, 11(3), 253–388. https://doi.org/10.1016/0883-0355(87)90001-2
- Marsh, H. W. (2007). Students’ evaluations of university teaching: A multidimensional perspective. In R. P. Perry & J. C. Smart (Eds.), The scholarship of teaching and learning in higher education: An evidence-based perspective (pp. 319–384). Springer. https://doi.org/10.1007/1-4020-5742-3_8
- Paas, F., Tuovinen, J. E., Tabbers, H., & Van Gerven, P. W. (2003). Cognitive load measurement as a means to advance cognitive load theory. Educational Psychologist, 38(1), 63–71. https://doi.org/10.1207/S15326985EP3801_8
- Rindermann, H. (1996). Untersuchungen zur Validität von Studentenevaluationen an Universitäten. Roderer.
- Rindermann, H. (2001). Lehrevaluation: Einführung und Bestandsaufnahme zu Forschung und Praxis. Empirische Pädagogik.
- Staufenbiel, T. (2000). Fragebogen zur Evaluation von Lehrveranstaltungen durch Studierende (FEVOR). Diagnostica, 46(4), 169–181. https://doi.org/10.1026//0012-1924.46.4.169
- Streiner, D. L. (2003). Starting at the beginning: An introduction to coefficient alpha and internal consistency. Journal of Personality Assessment, 80(1), 99–103. https://doi.org/10.1207/S15327752JPA8001_18
- Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257–285. https://doi.org/10.1207/s15516709cog1202_4
- Thielsch, M. T., Hirschfeld, G., & Haaser, I. (2008). Münsteraner Fragebogen zur Evaluation von Seminaren (MFE-S). In ZIS – GESIS Leibniz-Institut für Sozialwissenschaften. https://doi.org/10.6102/zis53
- Zimmerman, B. J. (2000). Attaining self-regulation: A social cognitive perspective. In M. Boekaerts, P. R. Pintrich, & M. Zeidner (Eds.), Handbook of self-regulation (pp. 13–39). Academic Press. https://doi.org/10.1016/B978-012109890-2/50031-7
13. Items of the Scale
The following are the 14 complete items of the Münster Questionnaire for the Evaluation of Seminars (MFE-S) as presented in the original German instrument, accompanied by English academic translations.
Instructions (Instruktion)
The seminar evaluation item batteries are presented without course-specific instructions because they are embedded within a centralized online evaluation system. General instructions regarding confidentiality, academic purpose, and digital navigation are provided upon first accessing the platform. Participant demographic information (age, gender, semester standing, and degree program) is queried at system login.
Rating Scale for Core Items (Items 2 to 12)
Responses to items 2 through 12 are recorded using a 7-point Likert scale defined by the following endpoints:
2
3
4
5
6
7 = völlig zutreffend (completely accurate)
Subscale 1: Didactics (Didaktik)
German: Der/Die Lehrende erläuterte schwierige Sachverhalte verständlich.
English translation: The instructor explained difficult concepts comprehensibly.
German: Man hat gemerkt, dass der Dozent/die Dozentin die Lehre für wichtig hält.
English translation: One could tell that the instructor considered teaching to be important.
German: Die Lehrmethoden waren zur Vermittlung des Stoffes gut geeignet.
English translation: The teaching methods were well-suited for imparting the course material.
Subscale 2: Difficulty (Schwierigkeit)
German: Der/Die Lehrende passte das Niveau des Seminars an den Wissensstand der Studierenden an.
English translation: The instructor adapted the level of the seminar to the students’ baseline knowledge.
German: Die Inhalte des Seminars waren zu schwierig für mich.
English translation: The seminar content was too difficult for me.
German: Die Lehrmethoden des Seminars waren wenig abwechslungsreich.
English translation: The teaching methods of the seminar lacked variety.
Subscale 3: Self-Study (Selbststudium)
German: Es wurden ausreichend Materialien (z.B. Literaturangaben, Skript) zur Vertiefung des Stoffes angeboten.
English translation: Sufficient materials (e.g., literature references, course reader) were provided to deepen the course content.
German: Es wurden hilfreiche Materialien (z.B. Literaturangaben, Skript) zur Vertiefung des Stoffes angeboten.
English translation: Helpful materials (e.g., literature references, course reader) were provided to deepen the course content.
German: Ich habe die Seminarsitzungen regelmäßig vorbereitet (z.B. durch das Lesen von Literatur oder die Bearbeitung von Hausaufgaben).
English translation: I prepared regularly for the seminar sessions (e.g., by reading literature or completing homework assignments).
Supplementary Diagnostic and Administrative Items
German: Was war Ihr HAUPTGRUND für den Besuch der Veranstaltung?
English translation: What was your PRIMARY REASON for attending this course?
• Interesse (Personal / Subject interest)
• Sonstiges (Other reasons)
German: Die Thematik hat mich schon vor der Veranstaltung sehr interessiert.
English translation: I was already very interested in the topic before the course began.
(Rated on the 1–7 Likert scale)
German: Ich habe in der Veranstaltung inhaltlich viel gelernt.
English translation: I learned a lot of content in this course.
(Rated on the 1–7 Likert scale)
German: Im Punktesystem der gymnasialen Oberstufe (0 [ungenügend] bis 15 [sehr gut +]) bewerte ich die Veranstaltung insgesamt mit folgender Punktzahl: ___
English translation: On the point scale of the upper secondary school (gymnasiale Oberstufe: 0 [insufficient/failing] to 15 [very good + / outstanding]), I rate the course overall with the following score: ___
(Open integer input from 0 to 15)
German: Anmerkungen für den/die Lehrende/n: Was hat Ihnen besonders gut an dieser Veranstaltung gefallen? Haben Sie Vorschläge für Veränderungen? Sonstige Anmerkungen:
English translation: Comments for the instructor: What did you like particularly well about this course? Do you have suggestions for changes? Other comments:
(Open text field for qualitative feedback)