Educational PsychologyHigher Education EvaluationPsychometrics

Münster Questionnaire for the Evaluation of Seminars – Revised (MFE-Sr)

A psychometric review and documentation of the Münster Questionnaire for the Evaluation of Seminars – Revised (MFE-Sr), covering construct definition, four-factor CFA structure, reliability, validity, and complete items.

memjavad
PUBLISHED
Scientifically Reviewed · Dr. Marwa Abd-Alazim · September 30, 2026
Medically & Scientifically Reviewed Verified: September 30, 2026
Dr. Marwa Abd-Alazim Ph.D.
Professor of Psychology • University of Kerbala
Review Criteria & Clinical Standards

This content undergoes rigorous scientific peer-review and medical editorial standards at Arab Psychology Network to ensure clinical accuracy, validity, and compliance with evidence-based guidelines from leading psychological and healthcare authorities (APA / WHO).

Abstract

The Münster Questionnaire for the Evaluation of Seminars – Revised (German: Münsteraner Fragebogen zur Evaluation von Seminaren – revidiert, abbreviated as MFE-Sr) is a specialized psychometric assessment tool developed at the University of Münster (Westfälische Wilhelms-Universität Münster) designed for student evaluations of educational quality within higher education seminar settings. Stemming from the modular Münster Evaluation System (Münsteraner Fragebogen zur Evaluation, MFE), the MFE-Sr was engineered to balance psychometric rigor with extreme methodological economy, providing actionable, multidimensional feedback while minimizing survey fatigue during high-stakes end-of-semester examination periods. The core psychometric core of the instrument comprises 15 primary items organized into four correlated first-order dimensions: Instructor & Didactics (Dozent & Didaktik; 6 items), Excessive Demand / Overload (Überforderung; 3 items), Participants / Peer Interaction (Teilnehmer; 3 items), and Instructional Materials (Materialien; 3 items). These core evaluative items are presented alongside 13 contextual and diagnostic items (yielding 28 total items) measuring attendance, prep time, curricular motives, spatial/environmental conditions, learning gain, and open-ended feedback.

Psychometric validation using confirmatory factor analysis (CFA) on a curated sample of N = 657 university students confirmed the hypothesized four-dimensional factor structure with robust goodness-of-fit indices: χ² = 301.0 (df = 84), Tucker-Lewis Index (TLI) = .95, Comparative Fit Index (CFI) = .96, and Root Mean Square Error of Approximation (RMSEA) = .06. The internal consistency coefficients (Cronbach’s α) demonstrate solid reliability across all subscales despite their brevity: Instructor & Didactics (α = .89), Excessive Demand (α = .77), Participants (α = .82), and Materials (α = .83). Criterion and construct validity are supported by substantial correlations with overall course satisfaction (ranging up to r = .89 for Instructor & Didactics) and significant discrimination across diverse course types (η² = .33). Core evaluative items employ a 7-point Likert response scale ranging from 1 (“strongly disagree”) to 7 (“strongly agree”), supplemented by a distinct “cannot be answered sensibly” option. The MFE-Sr represents a benchmark in higher education quality assurance, offering evidence-based diagnostic capability tailored for modern web-based institutional evaluation platforms.

Keywords

MFE-Sr, student evaluation of teaching, teaching quality, higher education quality assurance, psychometrics, confirmatory factor analysis, instructional design, academic overload, didactic competence, seminar evaluation

Authors

The Münster Questionnaire for the Evaluation of Seminars – Revised (MFE-Sr) was developed, standardized, and validated by academic researchers affiliated with the Institute of Psychology at the Westfälische Wilhelms-Universität Münster and partner institutions in Germany:

  • Dr. Meinald Thielsch, Dipl.-Psych. — Westfälische Wilhelms-Universität Münster, Psychologisches Institut 1, Fliednerstr. 21, 48149 Münster, Germany. E-Mail: [email protected]. Dr. Thielsch has spearheaded the institutional implementation, psychometric tracking, and methodological advancement of online teaching evaluations and web-based psychological research methodologies.
  • Dr. Gerrit Hirschfeld, Dipl.-Psych. — Vodafone Stiftungsinstitut und Lehrstuhl für Kinderschmerztherapie und Pädiatrische Palliativmedizin, Vestische Kinder- und Jugendklinik Datteln, Private Universität Witten/Herdecke, Dr.-Friedrich-Steiner-Str. 5, 45711 Datteln, Germany. E-Mail: [email protected]. Dr. Hirschfeld contributed extensive expertise in structural equation modeling, latent variable modeling, and quantitative scale development.

Purpose

The overarching purpose of the Münster Questionnaire for the Evaluation of Seminars – Revised (MFE-Sr) is to provide higher education institutions, department chairs, and university instructors with a psychometrically validated, standardized, and highly efficient instrument to assess instructional quality in interactive academic seminars. Student evaluations of educational quality (frequently abbreviated as SET or SEEQ) have served as an indispensable pillar of university quality management for over half a century (Schmidt & Loßnitzer, 2010). However, traditional evaluation instruments developed during the pen-and-paper era were often cumbersome, frequently consisting of 40 to 80 items. In the contemporary European higher education landscape shaped by the Bologna Process, academic semesters culminate in congested testing windows characterized by severe student exam stress and administrative overload (Bechler & Thielsch, 2012). Long questionnaires distributed under these conditions inevitably trigger survey fatigue, elevated non-response rates, and indiscriminate straight-lining or acquiescence bias.

The MFE-Sr was specifically engineered to overcome these systemic barriers by offering a hyper-economical core module tailored for web-based digital administration (Göritz et al., 2005; Haaser et al., 2007). In higher education, seminars possess fundamentally distinct pedagogical dynamics compared to mass lectures: while lectures prioritize unidirectional knowledge dissemination, seminars depend heavily on student-led presentations, reciprocal group discourse, preparation compliance, collaborative inquiry, and critical literature analysis. Consequently, evaluating seminars requires an instrument that does not solely evaluate the lecturer in isolation, but instead captures the multifaceted social and academic ecosystem of the classroom. The MFE-Sr fulfills three essential organizational objectives, as conceptualized in higher education quality research (Souvignier & Gold, 2002):

  • Formative Pedagogical Feedback: Delivering differentiated, diagnostic data to instructors regarding their pacing, clarity, thematic organization, and utilization of instructional media, thereby highlighting actionable pedagogical strengths and deficits.
  • Administrative and Institutional Steering: Providing department chairs, deans of academic affairs, and accreditation bodies with reliable, aggregated quality indicators that can inform curriculum refinement, teaching awards, continuing faculty education, and performance-based resource allocation.
  • Empirical Higher Education Research: Generating standardized, methodologically robust datasets to investigate systemic determinants of instructional effectiveness, such as the interplay between student workload, baseline course motivation, physical learning environments, and perceived learning gains.

Furthermore, the MFE-Sr purposefully integrates diagnostic bias variables—including student course prioritization (first choice vs. compulsory enrollment), self-reported absenteeism, weekly preparation hours, and environmental adequacy (acoustics and room facilities). By isolating these extraneous situational and motivational influences, the scale enables instructors and administrators to interpret student ratings within their proper ecological context.

Psychological Construct

The construct assessed by the MFE-Sr is multidimensional instructional quality in academic seminars. Rather than treating teaching quality as a monolithic or unifactorial attribute, modern psychometrics recognizes that classroom learning environments are composed of distinct, interacting pedagogical and psychosocial subsystems. The core psychometric architecture of the MFE-Sr comprises 15 items mapping onto four correlated first-order dimensions, complemented by diagnostic context variables:

1. Instructor & Didactics (Dozent & Didaktik)

The Instructor & Didactics subscale measures the pedagogical competence, communicative effectiveness, and structural organization exhibited by the instructor. In the setting of an academic seminar, didactic effectiveness does not merely entail clear oratory; it demands the skillful moderation of intellectual discourse, responsive engagement with student inquiries, cognitive scaffolding, and the ability to maintain conceptual coherence. This 6-item dimension assesses whether the instructor provided a structured overview of the disciplinary domain (Item 8), employed clarifying real-world examples (Item 9), responded constructively and respectfully to student questions and comments (Item 10), presented content in an intellectually stimulating manner (Item 11), maintained a transparent conceptual syllabus/outline throughout the semester (Item 12), and managed instructional time efficiently (Item 13). High scores on this subscale reflect superior pedagogical expertise, cognitive clarity, and interactive responsiveness.

2. Excessive Demand / Overload (Überforderung)

The Excessive Demand subscale quantifies student cognitive and temporal strain. Grounded in Cognitive Load Theory, academic learning fails when instructional demands exceed students’ working memory capacity or available study resources. This 3-item subscale captures subjective overburdening across three specific vectors: conceptual difficulty (Item 14: seminar content being too complex or difficult), instructional pacing (Item 15: the pace of knowledge transmission being excessively rapid), and out-of-class workload (Item 16: time investment required exceeding student coping capacity). In contrast to the other subscales, low-to-moderate values on this scale indicate healthy cognitive challenge, whereas elevated scores signal pathological cognitive overload, pacing mismatch, or excessive curricular demands.

3. Participants / Peer Interaction (Teilnehmer)

The Participants subscale reflects the academic engagement, preparation culture, and active participation of the student cohort. Unlike frontal lectures, the pedagogical success of a seminar is deeply co-constructed by the peer group. Even the most gifted instructor cannot compensate for a seminar where students attend unprepared or disengaged. This 3-item dimension operationalizes the social climate of learning by evaluating whether peers were adequately prepared for individual sessions (Item 17), actively contributed to seminar discussions (Item 18), and followed proceedings with attention and intellectual curiosity (Item 19). High scores indicate a vibrant, collaborative, and academically rigorous peer culture.

4. Instructional Materials (Materialien)

The Instructional Materials subscale evaluates the pedagogical utility, visual quality, and clarity of the physical and digital media integrated into the seminar. Grounded in multimodal learning theory, high-quality artifacts reduce extraneous cognitive load and reinforce deeper conceptual integration. This 3-item scale evaluates whether in-class media (presentation slides, instructional video clips, diagrams, schematic sketches) meaningfully contributed to comprehension (Item 20), whether the intrinsic visual and technical quality of these media was high (Item 21), and whether supplementary materials (e.g., reading packets, handouts, digital repository texts) met high academic standards (Item 22).

Contextual, Environmental, and Diagnostic Bias Variables

Beyond these four core latent factors, the MFE-Sr embeds 13 specialized operational items. Items 1 to 4 assess attendance fidelity, weekly study hours, course choice priority, and specific enrollment motives (interest, time constraints, mandatory course status, lack of alternatives). Items 5 to 7 evaluate structural prerequisites: spatial suitability (seminar room, breakout rooms), acoustic intelligibility, and schedule compatibility. Items 23 and 24 track the specific categories of supplementary materials utilized and perceived material volume. Finally, Items 25 to 28 capture summative outcome indicators: perceived subjective learning gain (binary yes/no), course recommendation willingness (binary yes/no), an overall holistic grade mapped to the German upper-secondary 0–15 points grading scale, and qualitative open-ended narrative feedback.

Theoretical Framework

The conceptual architecture of the MFE-Sr is firmly anchored in the Multimodal Conditional Model of Instructional Success (Multimodales Bedingungsmodell des Lehrerfolgs), formulated and expanded by educational psychologist Heiner Rindermann (1996, 2003, 2009). Rindermann’s model synthesizes decades of instructional effectiveness research, positing that academic learning outcomes are not determined by instructor behaviors alone, but emerge from complex, dynamic interactions among four primary clusters of determinants:

  1. Instructor Characteristics and Behaviors (Input): Teaching competencies, instructional clarity, communicative rapport, pacing, organization, and cognitive activation techniques.
  2. Student Characteristics (Input & Process): Cognitive baseline prerequisites, domain-specific prior knowledge, intrinsic interest, enrollment motivation, time investment, preparation discipline, and active classroom participation.
  3. Contextual and Structural Parameters (Frame Conditions): Group size, curricular status (compulsory vs. elective), spatial acoustics, physical room amenities, time scheduling, and digital learning infrastructure.
  4. Didactic Media and Content (Artifacts & Curricula): Quality of reading lists, visual slide design, task design, and conceptual difficulty of the syllabus.

By mapping its measurement dimensions directly onto this systemic model, the MFE-Sr achieves robust ecological validity. Traditional single-score evaluations committed the fundamental attribution error of educational measurement: assigning total responsibility for course success or failure exclusively to the instructor. The MFE-Sr remedies this fallacy by decomposing the seminar experience into instructor-led didactics, student-driven engagement, material utility, and cognitive overload, while simultaneously controlling for contextual confounders (e.g., room acoustics, compulsory course assignment). Furthermore, the theoretical grounding draws heavily upon Self-Determination Theory (Deci & Ryan, 2000), which highlights that classroom autonomy, perceived competence, and peer relatedness drive sustained learning motivation. In an academic seminar, the interactive dynamic captured by the Participants subscale reflects social relatedness and collective academic competence, which directly modulates instructional success.

Validity

The construct, convergent, discriminant, and criterion validity of the MFE-Sr have been empirically substantiated through multi-wave psychometric investigations conducted at the University of Münster (Hirschfeld & Thielsch, 2009; Thielsch & Weltzin, 2012):

Content Validity

The content validity of the instrument originates from a multi-stage iterative item generation protocol (Grabbe, 2003; Haaser, 2006). Based on systematic reviews of higher education literature regarding the observable characteristics of high-quality university seminars, initial pools of 17 and 14 items were refined, modified, and expanded to represent all essential facets of seminar quality without redundancy. Aligning with Rindermann’s multimodal framework, the primary operational focus remains on instructor behavioral didactics, while giving proportional representation to participant dynamics, cognitive overload, and instructional media.

Convergent and Criterion Validity

Convergent validity was evaluated at the course level across an aggregated dataset of 93 distinct seminar courses. Scale aggregate scores were correlated with students’ holistic summative evaluation of the overall seminar (measured via an open-ended 0–15 point academic grading metric). The observed correlations demonstrated robust, theoretically coherent relationships:

  • Instructor & Didactics correlated strongly with overall seminar rating: r = .89 (p < .001), corroborating the theoretical premise that instructor pedagogical delivery represents the primary determinant of perceived instructional quality.
  • Instructional Materials demonstrated a strong positive association with overall evaluation: r = .81 (p < .001).
  • Participants exhibited an equally pronounced positive correlation: r = .81 (p < .001), highlighting that peer engagement strongly influences collective course satisfaction.
  • Excessive Demand (Overload) demonstrated a moderate, statistically significant negative correlation: r = -.33 (p < .01), illustrating that disproportionate pacing, conceptual obscurity, and excessive time demands undermine student course appraisals.

Discriminant and Divergent Validity

Divergent validity was examined by correlating MFE-Sr subscales with potential non-instructional biasing factors, particularly students’ initial enrollment preference (first-choice vs. second-choice vs. third-choice course allocation). Spearman rank correlations between enrollment priority and the subscales Excessive Demand, Participants, and Instructional Materials were non-significant, demonstrating that pre-existing course allocation priorities do not artifactually bias these evaluative dimensions. Although a statistically significant rank correlation emerged between course preference and Instructor & Didactics (r = -.13, p < .05), the effect size is negligible (< 2% explained variance), establishing that ratings primarily reflect actual in-class didactic performance rather than prior course selection bias.

Furthermore, discriminant validity across course formats was confirmed using multivariate analysis of variance (MANOVA), with course type as the independent factor and the four MFE-Sr subscales as dependent variables. The omnibus test demonstrated significant between-course discrimination: F(368) = 3.06, p < .01, η² = .33. This indicates that one-third of the total multivariate variance in questionnaire ratings is directly attributable to systematic differences across courses, confirming that the MFE-Sr sensitively detects true pedagogical variation.

Reliability

The internal consistency of the MFE-Sr was evaluated using Cronbach’s α based on the standard validation sample of N = 657 completed, cleaned seminar evaluations collected during the winter semester 2009/2010. Despite the deliberate economy of the subscales (consisting of only 3 to 6 items each), all dimensions exceed conventional psychometric benchmarks for group-level diagnostic instruments:

  • Instructor & Didactics (6 items): α = .89. Corrected item-total correlations (T) range from .64 to .79, reflecting high scale homogeneity. The alpha coefficient remains stable between .86 and .89 if any single item is deleted.
  • Excessive Demand / Overload (3 items): α = .77. Corrected item-total correlations range from .48 to .69. Item 14 (content difficulty, T = .68) and Item 15 (pacing, T = .69) exhibit high consistency; Item 16 (time burden, T = .48) exhibits a lower item-total correlation, reflecting the reality that out-of-class study time is co-determined by individual external obligations. Deleting Item 16 would raise alpha to .86, but it is retained deliberately for its diagnostic utility to instructors.
  • Participants (3 items): α = .82. Corrected item-total correlations range from .58 to .74, demonstrating solid internal coherence regarding the peer learning climate.
  • Instructional Materials (3 items): α = .83. Corrected item-total correlations range from .57 to .75, indicating unified appraisal of physical, visual, and supplementary pedagogical resources.

These reliability values correspond favorably with much longer, established German instruments such as the FEVOR (Staufenbiel, 2000), HILVE (Rindermann, 2009), KIEL (Gediga et al., 2000), and TRIL (Gollwitzer & Schlotz, 2003), while requiring a fraction of respondent completion time.

Factor Analysis

The latent structure of the MFE-Sr was verified using Confirmatory Factor Analysis (CFA) executed in AMOS via Maximum Likelihood (ML) estimation on the N = 657 validation dataset. The hypothesized structural model posited four correlated first-order latent factors accounting for the covariance among the 15 core items.

Goodness-of-Fit Indices

The four-factor measurement model achieved an acceptable-to-good empirical fit according to established methodological standards (Hu & Bentler, 1999):

  • Chi-Square (χ²): 301.0 (df = 84, p < .001)
  • χ² / df Ratio: 3.58
  • Tucker-Lewis Index (TLI): .95
  • Comparative Fit Index (CFI): .96
  • Root Mean Square Error of Approximation (RMSEA): .06 (90% CI [.053, .068])

Standardized Factor Loadings and Item Statistics

The standardized factor loadings (λ), item means (M), standard deviations (SD), corrected item-total correlations (T), and Cronbach’s alpha if item deleted (αdel) are summarized below:

Subscale & Item Number M SD T Factor Loading (λ) α if Deleted
Instructor & Didactics (α = .89)
Item 8 (Overview of subject area) 5.94 1.18 .74 .79 .87
Item 9 (Use of explanatory examples) 5.85 1.28 .77 .82 .87
Item 10 (Responsiveness to student questions) 6.09 1.29 .71 .76 .87
Item 11 (Interesting presentation) 5.69 1.41 .79 .86 .86
Item 12 (Traceable structure/outline) 5.87 1.31 .65 .68 .88
Item 13 (Effective time management) 5.71 1.38 .64 .67 .89
Excessive Demand / Overload (α = .77)
Item 14 (Content too difficult) 2.22 1.27 .68 .86 .62
Item 15 (Pace too fast) 2.19 1.33 .69 .87 .60
Item 16 (Time expenditure overtaxing) 2.67 1.61 .48 .52 .86
Participants (α = .82)
Item 17 (Preparation of participants) 5.49 1.26 .58 .61 .85
Item 18 (Active contributions of peers) 5.58 1.18 .74 .82 .68
Item 19 (Attentive/interested participation) 5.61 1.20 .71 .90 .71
Instructional Materials (α = .83)
Item 20 (Media aided comprehension) 6.02 1.08 .75 .90 .70
Item 21 (Visual/technical media quality) 5.96 1.11 .75 .88 .70
Item 22 [listed in source table as Item 23] (Quality of additional materials) 5.66 1.25 .57 .62 .88

Inter-Factor Correlations and Descriptive Distributions

The inter-factor correlations (N = 657) illustrate that Instructor & Didactics is strongly and positively correlated with Instructional Materials (r = .76) and Participants (r = .66). The Participants and Instructional Materials subscales correlate at r = .65. In contrast, Excessive Demand correlates negatively with all three supportive learning dimensions: Instructor & Didactics (r = -.23), Participants (r = -.22), and Materials (r = -.19).

Descriptive statistics reveal substantial negative skewness across the positive dimensions (Instructor & Didactics: Median = 6.17, Mean = 5.86, SD = 1.06, Skewness = -1.58, Kurtosis = 3.10; Materials: Median = 6.00, Mean = 5.88, SD = 0.99, Skewness = -1.64, Kurtosis = 4.36; Participants: Median = 5.67, Mean = 5.56, SD = 1.04, Skewness = -0.96, Kurtosis = 1.28). In contrast, Excessive Demand displays positive skewness (Median = 2.00, Mean = 2.36, SD = 1.16, Skewness = 1.09, Kurtosis = 1.06), confirming that students generally experience low-to-moderate levels of overload in standard university seminar curricula.

Instrument / Measurement Tool

The MFE-Sr is structured as a multidimensional, standardized educational evaluation battery designed primarily for online computer-assisted web interviewing (CAWI) or mixed-mode administration.

  • Instrument Structure:
    • Core Psychometric Items (Items 8–22): 15 rating statements measuring four latent factors: Instructor & Didactics (Items 8–13), Excessive Demand (Items 14–16), Participants (Items 17–19), and Instructional Materials (Items 20–22).
    • Diagnostic Context and Control Items (Items 1–7): Single items capturing self-reported attendance (Item 1), weekly preparation/follow-up study hours (Item 2), course enrollment priority (Item 3), enrollment motives (Item 4), spatial/physical room quality (Item 5), room acoustics (Item 6), and scheduling convenience (Item 7).
    • Resource and Outcome Items (Items 23–28): Specific materials utilized (Item 23), perceived volume of materials (Item 24), subjective learning gain (Item 25), course recommendation (Item 26), summative overall grade on a 0–15 point scale (Item 27), and qualitative text commentary (Item 28).
  • Response Scale for Core Evaluative Items (Items 8–22, plus Items 5–7):
    • A 7-point Likert agreement scale is utilized: 1 = “stimme gar nicht zu” (strongly disagree), 2 = “stimme nicht zu” (disagree), 3 = “stimme eher nicht zu” (somewhat disagree), 4 = “neutral” (neutral), 5 = “stimme eher zu” (somewhat agree), 6 = “stimme zu” (agree), and 7 = “stimme vollkommen zu” (strongly agree).
    • An explicit non-substantive response option is provided: “nicht sinnvoll beantwortbar” (cannot be answered sensibly / not applicable).
  • Administration Rules & Multi-Instructor Delivery:
    • The instrument is administered without a separate instructional preamble; general instructions regarding data usage, confidentiality, and technical requirements are presented upon initial system login.
    • If a seminar is co-taught by multiple instructors, the Instructor & Didactics subscale (Items 8–13) is presented repeatedly and evaluated separately for each individual instructor. All remaining subscales and context items are answered once relative to the overall course.
    • Voluntary self-exclusion protocol: A completed evaluation dataset is only incorporated into departmental statistical analyses if the student respondent explicitly grants permission via a final confirmation prompt at the conclusion of the survey (Thielsch & Weltzin, 2012).
  • Scoring and Aggregation Procedures:
    • Given the unidimensionality of the individual subscales, scale scores can be computed either as the arithmetic mean (recommended: range 1.00 to 7.00) or sum score across the constituent items.
    • Items marked with “nicht sinnvoll beantwortbar” or left blank are treated as missing data. Casewise or pairwise exclusion rules are applied depending on institutional reporting thresholds.
    • Interpretation directionality: For Instructor & Didactics, Participants, and Instructional Materials, higher scores reflect superior instructional quality and favorable conditions. For Excessive Demand (Items 14–16), lower to moderate values (e.g., ≤ 3.00) are targeted; high scores indicate excessive difficulty, frantic pacing, or unsustainable student time burdens.

Permissions & Fee and Test Year

The Münster Questionnaire for the Evaluation of Seminars – Revised (MFE-Sr) was developed and validated in its revised format between 2009 and 2010 at the Westfälische Wilhelms-Universität Münster (Psychological Institute I), following earlier foundational iterations dating back to 2002/2003 (Grabbe, 2003; Haaser, 2006; Hirschfeld & Thielsch, 2009). The questionnaire was released into the academic public domain for scientific research, academic teaching evaluation, and institutional quality assurance purposes.

There are no licensing fees or commercial purchase requirements associated with utilizing the MFE-Sr for institutional quality assurance or non-commercial scholarly research. Higher education institutions and academic investigators are permitted to implement the items within web-based learning management systems (e.g., Moodle, ILIAS, EvaSys) or paper-pencil formats, provided proper academic attribution is maintained. Inquiries concerning system integration, normative comparative data benchmarks, or collaborative research should be directed to Dr. Meinald Thielsch at [email protected].

References

  • Bechler, S., & Thielsch, M. T. (2012). Prüfungsangst und Studienzufriedenheit in Bachelor- und Diplomstudiengängen [Test anxiety and study satisfaction in bachelor and diploma programs]. Zeitschrift für Pädagogische Psychologie, 26(4), 269–279. https://doi.org/10.1024/1010-0652/a000078
  • Deci, E. L., & Ryan, R. M. (2000). The “what” and “why” of goal pursuits: Human needs and the self-determination of behavior. Psychological Inquiry, 11(4), 227–268. https://doi.org/10.1207/S15327965PLI1104_01
  • Gediga, G., Schömann, K. M., & Hamborg, K. C. (2000). Der Kieler Fragebogen zur Vorlesungsevaluation (KIEL) [The Kiel questionnaire for lecture evaluation]. Zeitschrift für Pädagogische Psychologie, 14(4), 196–206.
  • Gollwitzer, M., & Schlotz, W. (2003). Das Trierer Inventar zur Lehrevaluation (TRIL) [The Trier inventory for teaching evaluation]. Diagnostica, 49(1), 34–44. https://doi.org/10.1026//0012-1924.49.1.34
  • Göritz, A. S., Soucek, R., & Bacher, J. (2005). Online-Befragungen [Online surveys]. In Handbuch Methoden der empirischen Sozialforschung (pp. 523–536). VS Verlag für Sozialwissenschaften.
  • Grabbe, Y. (2003). Entwicklung und Erprobung eines Fragebogens zur Evaluation von Seminaren [Development and testing of a questionnaire for seminar evaluation] (Unpublished diploma thesis). Westfälische Wilhelms-Universität Münster, Germany.
  • Greenwald, A. G. (1997). Validity concerns and usefulness of student ratings of instruction. American Psychologist, 52(11), 1182–1186. https://doi.org/10.1037/0003-066X.52.11.1182
  • Haaser, P. (2006). Optimierung des Fragebogens zur Evaluation von Seminaren am Fachbereich Psychologie und Sportwissenschaft [Optimization of the seminar evaluation questionnaire at the Department of Psychology and Sports Science] (Unpublished diploma thesis). Westfälische Wilhelms-Universität Münster, Germany.
  • Haaser, P., Thielsch, M. T., & Moeck, B. (2007). Das Münsteraner Online-Evaluationssystem [The Münster online evaluation system]. In M. Krämer, S. Preiser, & K. Brusdeylins (Eds.), Psychologiedidaktik und Evaluation VI (pp. 371–378). Shaker Verlag.
  • Hirschfeld, G., & Thielsch, M. T. (2009). Konfirmatorische Prüfung des MFE-S [Confirmatory testing of the MFE-S] (Internal Research Report). Westfälische Wilhelms-Universität Münster, Germany.
  • Hu, L. T., & Bentler, P. M. (1999). Cutoff criteria for fit indexes in covariance structure analysis: Conventional criteria versus new alternatives. Structural Equation Modeling: A Multidisciplinary Journal, 6(1), 1–55. https://doi.org/10.1080/10705519909540118
  • Marsh, H. W. (1984). Students’ evaluations of university teaching: Dimensionality, reliability, validity, potential biases, and utility. Journal of Educational Psychology, 76(5), 707–754. https://doi.org/10.1037/0022-0663.76.5.707
  • Mutz, R. (2003). Die Güte von Studierendenurteilen über die Lehrqualität: Konfirmatorische Multitrait-Multimethod-Analysen zur Validität der studentischen Lehrveranstaltungsbewertung [The quality of student judgments on teaching quality: Confirmatory multitrait-multimethod analyses on the validity of student ratings of instruction]. Empirische Pädagogik, 17(3), 362–386.
  • Rindermann, H. (1996). Untersuchungen zur Validität von Studentenevaluationen [Studies on the validity of student evaluations]. Roderer.
  • Rindermann, H. (2003). Empfehlungen für die Evaluation universitärer Lehre [Recommendations for the evaluation of university teaching]. Empirische Pädagogik, 17(3), 434–467.
  • Rindermann, H. (2009). Lehrevaluation: Einführung und Überblick zu Forschung und Praxis [Teaching evaluation: Introduction and overview of research and practice] (2nd ed.). Verlag Empirische Pädagogik.
  • Schmidt, F. T., & Loßnitzer, T. (2010). Praxishandbuch Lehrevaluation [Practical handbook of teaching evaluation]. Verlag Empirische Pädagogik.
  • Souvignier, E., & Gold, A. (2002). Lehrevaluation: Feedback, Steuerung, Forschung [Teaching evaluation: Feedback, governance, research]. Zeitschrift für Pädagogische Psychologie, 16(3/4), 143–146. https://doi.org/10.1024//1010-0652.16.34.143
  • Staufenbiel, T. (2000). Fragebogen zur Evaluation von universitären Lehrveranstaltungen (FEVOR) [Questionnaire for the evaluation of university courses]. Diagnostica, 46(4), 169–181. https://doi.org/10.1026//0012-1924.46.4.169
  • Thielsch, M. T., & Weltzin, S. (2012). Online-Lehrevaluation: Praxisleitfaden für Hochschulen [Online teaching evaluation: Practical guide for universities]. Universitätsverlag Münster.
  • Thielsch, M. T., Hirschfeld, G., & Haaser, P. (2010). Rücklaufquoten und Akzeptanz webbasierter Lehrevaluationen an Hochschulen [Response rates and acceptance of web-based teaching evaluations in higher education]. In M. Krämer, S. Preiser, & K. Brusdeylins (Eds.), Psychologiedidaktik und Evaluation VIII (pp. 289–297). Shaker Verlag.

Items of the Scale

The authentic original German items of the Münster Questionnaire for the Evaluation of Seminars – Revised (MFE-Sr) are documented below verbatim along with their official response formats.

Antwortformat für die Kernskalen (Items 8–22 sowie Items 5–7)

Für Items 8-22 wird ein 7-stufiges Antwortformat mit den Optionen 1 = “stimme gar nicht zu”, 2 = “stimme nicht zu”, 3 = “stimme eher nicht zu”, 4 = “neutral”, 5 = “stimme eher zu”, 6 = “stimme zu” und 7 = “stimme vollkommen zu” verwendet. Zusätzlich steht die Antwortoption “nicht sinnvoll beantwortbar” zur Verfügung.


1. Items zu Dozent & Didaktik

  1. Item 8: Ich finde, das Seminar gab einen guten Überblick über das Themengebiet.
  2. Item 9: Der/Die Lehrende benutzte oft Beispiele, die zum Verständnis der Lehrinhalte beitrugen.
  3. Item 10: Ich finde, der/die Lehrende ging auf Fragen und Anregungen der Studierenden angemessen ein.
  4. Item 11: Der/Die Lehrende hat das Thema interessant aufgearbeitet.
  5. Item 12: Ich konnte im Verlauf des Seminars die Gliederung immer nachvollziehen.
  6. Item 13: Ich finde, der/die Lehrende teilte die zur Verfügung stehende Zeit gut ein.

2. Items zu Überforderung

  1. Item 14: Die Inhalte des Seminars waren zu schwierig für mich.
  2. Item 15: Das Tempo der Stoffvermittlung war für mich zu hoch.
  3. Item 16: Der mit dem Seminar verbundene Zeitaufwand hat mich überfordert.

3. Items zu Teilnehmer

  1. Item 17: Die meisten Teilnehmer waren gut auf die einzelnen Termine vorbereitet.
  2. Item 18: Die meisten Teilnehmer brachten sich aktiv ein.
  3. Item 19: Die meisten Teilnehmer verfolgten das Seminar aufmerksam und mit Interesse.

4. Items zu Materialien

  1. Item 20: Die in der Vorlesung verwendeten Medien (Folien, Filme, Skizzen, etc.) trugen zum Verständnis der Inhalte bei.
  2. Item 21: Die Qualität der in der Vorlesung verwendeten Medien (Folien, Filme, Skizzen, etc.) war gut.
  3. Item 22: Die Qualität der zusätzlichen Materialien war gut.

Zusätzlich vorgegebene, diagnostische und kontextuelle Items

  1. Item 1: Wie viele Sitzungen hast Du bei diesem Seminar gefehlt?
    Antwortoptionen: keine, eine, zwei, drei oder mehr Sitzungen
  2. Item 2: Wie viele Stunden hast Du das Seminar im Schnitt pro Woche vor- und nachbereitet?
    Antwortoption: Offenes Antwortfeld
  3. Item 3: Das Seminar war meine:
    Antwortoptionen: Erstwahl, Zweitwahl, Drittwahl oder geringer
  4. Item 4: Ich habe dieses spezielle Seminar gewählt (Mehrfachantworten möglich):
    Antwortoptionen: aus Zeitgründen, aus Interesse am Thema, wegen des/r Dozenten/in, kein alternativer Kurs, es waren bei der Wahl keine Informationen verfügbar, weil es eine Pflichtveranstaltung ist
  5. Item 5: Ich finde, die räumliche Ausstattung (Seminarraum, Gruppenräume, Experimentalräume) war angemessen.
    Antwortformat: Wie bei Items 8-22 (7-stufiges Antwortformat: 1 = “stimme gar nicht zu” bis 7 = “stimme vollkommen zu”, plus “nicht sinnvoll beantwortbar”)
  6. Item 6: Die Lautstärke war so, dass ich immer alles gut verstehen konnte.
    Antwortformat: Wie bei Items 8-22 (7-stufiges Antwortformat: 1 = “stimme gar nicht zu” bis 7 = “stimme vollkommen zu”, plus “nicht sinnvoll beantwortbar”)
  7. Item 7: Der Seminartermin passte gut in meine Zeitplanung.
    Antwortformat: Wie bei Items 8-22 (7-stufiges Antwortformat: 1 = “stimme gar nicht zu” bis 7 = “stimme vollkommen zu”, plus “nicht sinnvoll beantwortbar”)
  8. Item 23: Ich habe folgende Materialien zusätzlich zur Veranstaltung benutzt (Mehrfachantworten möglich):
    Antwortoptionen: keine, Folien, Skript, Literaturangaben, Webseite des Dozenten/der Veranstaltung, andere Webseiten, Handout, Sonstiges
  9. Item 24: Ich fand die Menge des Materials, das zu dieser Veranstaltung zur Verfügung gestellt wurde, war…
    Antwortoptionen: zu gering, angemessen, zu groß, nicht sinnvoll beantwortbar
  10. Item 25: Ich habe in der Veranstaltung viel gelernt.
    Antwortoptionen: Ja, Nein
  11. Item 26: Ich würde dieses Seminar anderen Studierenden weiterempfehlen.
    Antwortoptionen: Ja, Nein
  12. Item 27: Im Punktesystem der gymnasialen Oberstufe (0 [ungenügend] bis 15 [sehr gut +]) bewerte ich dieses Seminar mit folgender Punktzahl: ___
    Antwortoption: Offenes Antwortfeld (0 bis 15 Punkte)
  13. Item 28: Anmerkungen für den/die Lehrende/n (Vorschläge/Lob/konstruktive Kritik):
    Antwortoption: Offenes Antwortfeld
★

Rate This Scale

5.0 / 5 • 1 vote

Cite This Article

memjavad (2026, September 30). Münster Questionnaire for the Evaluation of Seminars – Revised (MFE-Sr). PSYCHOLOGICAL DATABASE. https://en.arabpsychology.com/scales/munster-questionnaire-for-the-evaluation-of-seminars-revised-mfe-sr/
memjavad. “Münster Questionnaire for the Evaluation of Seminars – Revised (MFE-Sr).” PSYCHOLOGICAL DATABASE, 30 September 2026, https://en.arabpsychology.com/scales/munster-questionnaire-for-the-evaluation-of-seminars-revised-mfe-sr/.
memjavad. “Münster Questionnaire for the Evaluation of Seminars – Revised (MFE-Sr).” PSYCHOLOGICAL DATABASE. September 30, 2026. https://en.arabpsychology.com/scales/munster-questionnaire-for-the-evaluation-of-seminars-revised-mfe-sr/.