Educational PsychologyEvaluation InstrumentsPsychometrics

Münster Questionnaire for the Evaluation of Lectures – Revised (MFE-Vr)

A psychometric review of the Münster Questionnaire for the Evaluation of Lectures – Revised (MFE-Vr), developed by Meinald T. Thielsch and Gerrit Hirschfeld for higher education course evaluations.

memjavad
PUBLISHED
Scientifically Reviewed · Dr. Marwa Abd-Alazim · September 30, 2026
Medically & Scientifically Reviewed Verified: September 30, 2026
Dr. Marwa Abd-Alazim Ph.D.
Professor of Psychology • University of Kerbala
Review Criteria & Clinical Standards

This content undergoes rigorous scientific peer-review and medical editorial standards at Arab Psychology Network to ensure clinical accuracy, validity, and compliance with evidence-based guidelines from leading psychological and healthcare authorities (APA / WHO).

Abstract

The Münster Questionnaire for the Evaluation of Lectures – Revised (German: Münsteraner Fragebogen zur Evaluation von Vorlesungen – revidiert, abbreviated as MFE-Vr) is an economically constructed, psychometrically validated multidimensional survey instrument designed for the higher education sector. Developed within the Department of Psychology at the University of Münster (Westfälische Wilhelms-Universität Münster) by Meinald T. Thielsch, Gerrit Hirschfeld, and colleagues, the MFE-Vr represents a standardized modular core system within the overarching Münster Questionnaire for Evaluation (MFE) framework. Built to support both diagnostic quality assurance and formative instructional feedback, the instrument measures students' perceptions of academic lectures while maintaining high administrative economy and low survey burden during peak examination periods.

The complete standardized inventory captures key evaluative facets across core pedagogical dimensions: instructional performance and didactic execution (Dozent & Didaktik), cognitive overload and pacing (Überforderung), instructional media and course materials (Materialien), as well as supplementary curricular, environmental, and behavioral contextual variables (e.g., student attendance, prior interest, acoustic suitability, room capacity, and global performance grading). The core psychometric scales utilize a 7-point Likert-type response scale ranging from 1 ("stimme gar nicht zu" / strongly disagree) to 7 ("stimme vollkommen zu" / strongly agree), supplemented by a distinct "nicht sinnvoll beantwortbar" (cannot be sensibly answered) residual category to safeguard against forced invalid responses.

Empirical validation studies involving university lecture cohorts demonstrate robust psychometric properties. Confirmatory factor analyses (CFA) verify that the multidimensional construct structure exhibits sound global fit indices ($\chi^2 = 187.0$, $df = 51$, $ ext{CFI} = .96$,$ ext{TLI} = .95$,$ ext{RMSEA} = .08$). Internal consistency estimates across subscales demonstrate good to excellent reliability, with Cronbach's alpha coefficients ranging from $lpha = .81$ to $lpha = .93$. Convergent validity is evidenced by strong correlations with global lecture grade benchmarks ($r = .90$ for the instructor/didactics dimension), while divergent validity is confirmed via negligible associations with situational bias variables ($r < .30$). Criterion-related discriminant power is underscored by large effect sizes separating high-performing from underperforming lecture courses, rendering the MFE-Vr an indispensable instrument for tertiary quality management, pedagogical development, and higher education research.

Keywords

Student Evaluation of Educational Quality, Münster Questionnaire for Evaluation, MFE-Vr, Lecture Evaluation, Higher Education Quality Assurance, Didactic Competence, Cognitive Overload, Confirmatory Factor Analysis, Psychometrics, Academic Teaching Assessment

Authors

The Münster Questionnaire for the Evaluation of Lectures – Revised was developed, refined, and psychometrically standardized by researchers affiliated with the Westfälische Wilhelms-Universität Münster (University of Münster) and partner clinical/academic institutions in Germany:

  • Dr. Dipl.-Psych. Meinald T. Thielsch: Department of Psychology (Psychologisches Institut 1), Westfälische Wilhelms-Universität Münster, Fliednerstraße 21, 48149 Münster, Germany. Email: [email protected]. Specialized in organizational psychology, psychometrics, online assessment, human-computer interaction, and quality management in tertiary education.
  • Dr. Dipl.-Psych. Gerrit Hirschfeld: Vodafone Foundation Institute and Chair of Children's Pain Therapy and Paediatric Palliative Medicine, Vestische Children's and Youth Clinic Datteln, Witten/Herdecke University, Dr.-Friedrich-Steiner-Straße 5, 45711 Datteln, Germany. Email: [email protected]. Specialized in advanced quantitative methods, latent variable modeling, psychometric scale development, and pediatric clinical outcomes.
  • Collaborating Contributors & Precursor Developers: Earlier modular versions and developmental predecessors at the University of Münster were informed by psychometric foundational work led by Dipl.-Psych. C. Grabbe (2003) and Dipl.-Psych. M. Haaser (2006), alongside institutional contributions from T. Moeck, M. Weltzin, and C. Bechler.

Purpose

The primary purpose of the Münster Questionnaire for the Evaluation of Lectures – Revised (MFE-Vr) is to provide a brief, methodologically sound, and standardized diagnostic measurement tool capable of evaluating academic lectures in higher education institutions. Student evaluations of teaching (SET) have become a cornerstone of university governance, mandated institutional accreditation, internal quality assurance, and instructional enhancement. However, a persistent challenge in academic survey research is the trade-off between conceptual comprehensiveness and survey economy. Traditional German-language teaching evaluation batteries, such as the Heidelberger Inventar zur Lehrveranstaltungs-Evaluation (HILVE; Rindermann, 2009) or the Fragebogen zur Evaluation von Vorlesungen (FEVOR; Staufenbiel, 2000), provide rich multi-attribute profiles but often demand considerable administration time. When administered during the concluding weeks of a semester—a timeframe that coincides with heightened student stress and academic examination preparation (Bechler & Thielsch, 2012)—lengthy questionnaires frequently suffer from survey fatigue, central tendency response bias, careless responding, and declining response rates.

The MFE-Vr addresses these methodological challenges by establishing an optimized, short-form item battery specifically tailored for automated web-based platforms and hybrid (mixed-mode) survey routines. Specifically, the instrument fulfills three interconnected academic functions classified under the conceptual taxonomy articulated by Souvignier and Gold (2002):

  • Formative Feedback and Instructional Diagnosis: The scale delivers granular, actionable intelligence directly to university instructors regarding lecture clarity, rhetorical pacing, instructional media quality, and student cognitive overload. Because the subscale structure isolates distinct facets of teaching performance, faculty members can pinpoint precise areas requiring modification (e.g., slowing down thematic presentation or improving lecture slides) without confounding teaching style with external infrastructure deficits.
  • Administrative Steering and Institutional Benchmarking: The aggregated outcomes enable departmental leadership, deans of study, and quality management committees to monitor curriculum effectiveness across academic terms. Within the automated platform operated at the University of Münster, raw scores are dynamically calibrated against institutional benchmarks and historical normative distributions across differing course categories, providing objective metrics for teaching awards, accreditation reporting, and instructional resource allocation.
  • Educational Research and Bias Control: Beyond pedagogical diagnostics, the MFE-Vr systematically captures potential confounding variables (e.g., lecture hall acoustics, scheduling conflicts, physical seating constraints, and prior student motivation). This structural separation enables educational researchers and institutional analysts to disentangle valid instructional effects from situational or environmental biases that often plague course ratings.

Psychological Construct

The MFE-Vr measures academic instructional quality as a multidimensional latent construct. Grounded in educational psychology, cognitive load research, and instructional communication models, the instrument operationalizes effective lecturing through core primary subscales and supplementary contextual indicators:

1. Instructor and Didactics (Dozent & Didaktik)

This primary subscale captures the pedagogical execution, communicative clarity, rhetorical structuring, and interpersonal responsiveness of the lecturer. Rather than measuring a generic affective impression or instructor popularity, the six items evaluate overt behavioral indicators of instructional effectiveness:

  • Thematic Structuring and Orientation: Providing an overarching contextual framework (e.g., Item 8: "Ich finde, die Vorlesung gab einen guten Überblick über das Themengebiet") and transparent signposting that allows students to follow the logical progression throughout the semester (Item 12).
  • Exemplification and Cognitive Elaboration: The purposeful integration of concrete examples and real-world applications to bridge abstract theoretical concepts with empirical phenomena (Item 9: "Der/Die Lehrende benutzte oft Beispiele, die zum Verständnis der Lehrinhalte beitrugen").
  • Interactive Receptivity: Addressing audience queries, welcoming student contributions, and demonstrating intellectual accessibility within a large auditorium setting (Item 10: "Ich finde, der/die Lehrende ging auf Fragen und Anregungen der Studierenden angemessen ein").
  • Rhetorical Engagement and Time Management: Delivering subject matter in an engaging manner (Item 11) while adhering to temporal parameters and pacing allocations (Item 13).

2. Cognitive Overload (Überforderung)

Directly grounded in Cognitive Load Theory (Sweller, 1988; Paas et al., 2003), this three-item subscale measures the subjective strain experienced by students due to excessive cognitive demands, conceptual difficulty, or rapid instructional delivery. In higher education lectures, cognitive failure occurs when total cognitive load exceeds working memory capacity. The subscale evaluates:

  • Conceptual Intrinsic Difficulty: Perception that lecture material surpassed the learners' cognitive baseline (Item 14: "Die Inhalte der Vorlesung waren zu schwierig für mich").
  • Pacing and Presentation Rate: The speed at which novel instructional material is introduced, which dictates cognitive processing intervals (Item 15: "Das Tempo der Stoffvermittlung war für mich zu hoch").
  • Extracurricular Temporal Demands: The degree to which preparation, review, and follow-up activities overwhelmed the student's available study time (Item 16: "Der mit der Vorlesung verbundene Zeitaufwand hat mich überfordert").

Unlike didactic quality, where high scores reflect positive outcomes, optimal instructional conditions on the Überforderung scale are denoted by low to moderate mean values, indicating that students were challenged without experiencing cognitive exhaustion.

3. Instructional Media and Accompanying Materials (Materialien)

This three-item subscale evaluates the pedagogical utility, design quality, and comprehensibility of multimedia aids and auxiliary study resources. In accordance with Mayer's Cognitive Theory of Multimedia Learning, educational presentations should facilitate dual-channel auditory and visual processing without producing split-attention or coherence deficits. The scale examines:

  • Functional Explanatory Power: The extent to which visual aids, slides, diagrams, and illustrative video demonstrations directly clarify the conceptual content (Item 17: "Die in der Vorlesung verwendeten Medien trugen zum Verständnis der Inhalte bei").
  • Visual and Design Quality: Legibility, technical execution, layout consistency, and aesthetic clarity of presentation materials (Item 18).
  • Quality of Supplementary Study Resources: The comprehensiveness and clarity of distributed lecture notes, syllabi, reading packets, or online portal uploads designed for autonomous post-lecture review (Item 19: "Die Qualität der zusätzlichen Materialien war gut").

4. Supplementary Contextual and Methodological Bias Variables

The broader MFE-Vr battery incorporates auxiliary non-latent screening items capturing contextual framing conditions. These encompass physical learning environments (auditorium acoustics, seating availability, room temperature/size), timetable integration, student attendance frequency, prior intrinsic interest versus mandatory enrolment motivations, and global summary marks. Isolating these indicators prevents systemic bias from contaminating evaluations of instructor performance.

Theoretical Framework

The conceptual architecture of the MFE-Vr is anchored in contemporary models of higher education instruction, psychometric test design, and instructional communication. Historically, student evaluations of teaching have been subjected to rigorous methodological scrutiny regarding their construct validity, susceptibility to external confounds, and susceptibility to the "Dr. Fox effect"—the phenomenon where charismatic presentation masks deficient substantive content (Greenwald, 1997; Marsh, 1984, 2007; Rindermann, 1996, 2003, 2009).

The MFE-Vr relies fundamentally on the Multimodal Conditions Model of Instructional Success (Multimodales Bedingungsmodell des Lehrerfolgs; Rindermann, 1996, 2009). This systemic paradigm conceptualizes educational effectiveness in tertiary classrooms as the product of multiple interacting clusters:

  1. Instructor Input Variables: Knowledge base, didactic communication skills, rhetorical pacing, preparation, and empathy.
  2. Student Input Variables: Prior knowledge, cognitive capability, intrinsic motivation, self-regulated study habits, and attendance patterns.
  3. Contextual and Structural Parameters: Class size, mandatory versus elective status, room acoustics, technological infrastructure, and timetable suitability.
  4. Process Characteristics: Active cognitive elaboration, open communicative climate, mental strain, and comprehension of lecture structure.
  5. Instructional Outcomes: Short-term mastery, deep conceptual learning, student satisfaction, and long-term interest in the discipline.

Within this theoretical model, empirical investigations demonstrate that instructor behavioral actions (Verhaltensweisen des Dozenten) represent the single most potent conditional determinant of student learning gains (Rindermann, 2009). By prioritizing behavioral didactic markers rather than vague personality impressions, the MFE-Vr maintains close adherence to the actionable components of the model. Furthermore, the explicit incorporation of Sweller's Cognitive Load Theory ensures that instructional pacing and intrinsic challenge are treated not as peripheral noise, but as core process variables that directly predict academic achievement and instructional attrition.

Validity

The construct, convergent, discriminant, and criterion-related validity of the MFE-Vr have been demonstrated across empirical studies utilizing both exploratory and confirmatory modeling frameworks (Grabbe, 2003; Haaser, 2006; Hirschfeld & Thielsch, 2009; Thielsch & Weltzin, 2012).

Content and Face Validity

Because universal objective criteria for "ideal teaching" do not exist in absolute terms (Marsh, 1984), content validity in the MFE-Vr was established through systematic expert reviews, institutional Delphi consultations, and iterative reductions of an initial 29-item pool capturing eight theoretical teaching dimensions (Grabbe, 2003). Items with marginal communicative relevance or ambiguous didactic interpretation were phased out, retaining only behavioral indicators that align with documented predictors of instructional success.

Convergent and Criterion Validity

Convergent validity was evaluated by examining bivariate relationships between the core MFE-Vr subscales and independent global evaluations of the lecture course (overall grade using the 15-point German upper secondary grading scale; Item 24):

  • Dozent & Didaktik demonstrated a very strong positive association with global course ratings ($r = .90$, $p < .001$), confirming that didactic competence forms the dominant core of students' overarching quality appraisal.
  • Materialien showed a strong, statistically significant correlation with global performance ($r = .77$, $p < .001$).
  • Überforderung exhibited a moderate, statistically significant negative correlation ($r = -.29$, $p < .01$), corroborating theoretical assertions that excessive cognitive overload detracted from perceived lecture quality.

Discriminant and Divergent Validity

Divergent validity is demonstrated by negligible associations between the core didactic subscales and extrinsic environmental bias factors. Student ratings of whether the lecture hall was sufficiently large or whether the class schedule fit personal routines correlated at trivial levels ($r < .30$) with didactic effectiveness ratings, showing that physical lecture hall conditions do not artificially distort instructional scores.

Furthermore, strong criterion-related discriminant validity was verified by evaluating the instrument's capacity to differentiate between courses independently identified as top-performing versus lowest-performing cohorts based on final marks:

  • Dozent & Didaktik significantly differentiated the top and bottom lecture offerings: $F(1, 65) = 91.43$, $p < .001$, $\eta^2 = .59$, representing a massive discriminative effect.
  • Überforderung significantly differentiated best from worst lectures: $F(1, 65) = 11.85$, $p < .01$, $\eta^2 = .15$.
  • Materialien likewise revealed clear discriminative separation: $F(1, 65) = 62.23$, $p < .001$, $\eta^2 = .49$.

When evaluated across the entire distribution of all 26 evaluated lecture courses, the omnibus effects remained substantial and statistically significant across all dimensions ($F_{ ext{Didaktik}}(25, 403) = 17.53$, $p < .001$, $\eta^2 = .52$; $F_{ ext{Überforderung}}(25, 403) = 3.32$, $p < .001$, $\eta^2 = .17$; $F_{ ext{Materialien}}(25, 403) = 9.80$, $p < .001$, $\eta^2 = .38$), demonstrating the sensitivity of the instrument to between-course variations.

Reliability

The reliability of the MFE-Vr subscales has been corroborated across successive student cohorts at the University of Münster. In a representative psychometric investigation involving a cleaned validation cohort of $N = 429$ students evaluating lecture courses during the Winter Semester 2009/2010, the scale demonstrated high internal consistency using Cronbach's alpha ($lpha$):

  • Dozent & Didaktik (6 items): $lpha = .93$, indicating exceptional internal consistency. Corrected item-total correlations ($r_{it}$) across the six items ranged from $.76$ to $.86$, reflecting high item homogeneity without redundancy.
  • Überforderung (3 items): $lpha = .85$, indicating good internal consistency for an ultra-brief three-item scale. Corrected item-total correlations ranged between $.65$ and $.78$.
  • Materialien (3 items): $lpha = .81$, indicating sound reliability for a brief scale. Corrected item-total correlations ranged from $.49$ to $.76$.

These reliability parameters compare favorably to lengthy standardized inventories (such as FEVOR, $lpha = .80–.94$; HILVE, $lpha = .78–.92$; or TRIL, $lpha = .82–.93$), confirming that the brevity of the MFE-Vr does not undermine measurement precision. Inter-rater reliability aggregated at the course level shows substantial agreement among enrolled students, satisfying institutional criteria for public quality reporting.

Factor Analysis

The structural dimensionality of the MFE-Vr was verified through linear structural equation modeling and confirmatory factor analysis (CFA) using AMOS, employing maximum likelihood (ML) estimation on the Winter Semester 2009/2010 validation sample ($N = 429$). The hypothesized three-factor measurement model assigned the 12 core items to their respective theoretical latent dimensions: Dozent & Didaktik (Items 8–13), Überforderung (Items 14–16), and Materialien (Items 17–19).

Goodness-of-Fit Statistics

The three-factor model yielded satisfactory global fit indices adhering to conventional psychometric thresholds (Hu & Bentler, 1999):

  • $\chi^2 = 187.0$ with $df = 51$ ($p < .001$)
  • Comparative Fit Index ($ ext{CFI}$) =$.96$
  • Tucker-Lewis Index ($ ext{TLI}$) =$.95$
  • Root Mean Square Error of Approximation ($ ext{RMSEA}$) =$.080$ ($90% \text{ CI } [.068, .093]$)

Factor Loadings and Latent Intercorrelations

Standardized factor loadings ($lambda$) across all subscales demonstrated moderate to strong saturation on their targeted latent constructs:

  • Dozent & Didaktik: Item 8 ($lambda = .81$), Item 9 ($lambda = .88$), Item 10 ($lambda = .82$), Item 11 ($lambda = .90$), Item 12 ($lambda = .81$), Item 13 ($lambda = .78$).
  • Überforderung: Item 14 ($lambda = .80$), Item 15 ($lambda = .92$), Item 16 ($lambda = .71$).
  • Materialien: Item 17 ($lambda = .94$), Item 18 ($lambda = .90$), Item 19 ($lambda = .50$).

Latent factor intercorrelations substantiated theoretical assumptions regarding instructional constructs. The latent didactic dimension (Dozent & Didaktik) correlated positively with instructional materials (Materialien; $r = .87$, $p < .001$) and moderately negatively with student cognitive strain (Überforderung; $r = -.29$, $p < .01$). The correlation between cognitive strain and instructional materials was similarly inverse ($r = -.28$, $p < .01$). These interrelationships confirm that while instructional presentation and materials share substantial common variance, cognitive overload constitutes a distinct, dissociable latent factor.

Instrument / Measurement Tool

The Münster Questionnaire for the Evaluation of Lectures – Revised (MFE-Vr) is administered primarily as an automated, web-based survey module, with paper-and-pencil formats reserved for specialized mixed-mode applications. Its administrative specifications are summarized below:

  • Test Type: Multidimensional Student Evaluation of Teaching (SET) inventory / course rating scale.
  • Target Population: Undergraduate and graduate students attending university lectures across academic faculties.
  • Administration Modality: Online survey execution via secure PHP/MySQL platforms (unipark/custom portals) or physical optical mark reader (OMR) paper forms.
  • Item Inventory Composition: The full modular survey battery comprises 25 numbered items categorized into:
    • Standard Didactic & Structural Dimensions: 12 primary Likert-type items forming three validated psychometric subscales (Dozent & Didaktik, Überforderung, Materialien). In the broader expanded framework, items cover five thematic clusters of 5 items each (Didaktik, Anregung, Interaktion, Struktur, Rahmenbedingungen).
    • Contextual Bias & Behavioral Screeners: Items 1–3, 20–21 (attendance frequency, weekly preparation/follow-up study hours, attendance motivations, specific learning resources used, volume of course materials).
    • Environmental Baseline Screeners: Items 4–7 (hall seating adequacy, room acoustics, volume intelligibility, timetable feasibility).
    • Global & Qualitative Criteria: Items 22–25 (binary subjective learning gain, peer recommendation, 15-point overall performance mark, and qualitative open-ended text commentary).
  • Response Format: The primary psychometric items are rated on an authentic 7-point Likert-type agreement scale:
    • $1$ = "stimme gar nicht zu" (Strongly disagree)
    • $2$ = "stimme nicht zu" (Disagree)
    • $3$ = "stimme eher nicht zu" (Somewhat disagree)
    • $4$ = "neutral" (Neutral)
    • $5$ = "stimme eher zu" (Somewhat agree)
    • $6$ = "stimme zu" (Agree)
    • $7$ = "stimme vollkommen zu" (Strongly agree)
    • Residual Category = "nicht sinnvoll beantwortbar" (Cannot be sensibly answered / not applicable).
  • Scoring and Aggregation Rules:
    • Given the confirmed unidimensionality of the core subscales, scale scores can be computed as either unweighted sum scores or arithmetic item means ($M$).
    • For multi-instructor team-taught lectures, the Dozent & Didaktik block is presented iteratively for each instructor individually, whereas overall lecture items (Materialien, Überforderung, environment) are presented once.
    • Directionality: High scores on Dozent & Didaktik and Materialien reflect positive didactic quality. For Überforderung, lower to moderate scores are desirable, as high values indicate excessive cognitive strain.
    • Reverse-Coding Rules: In the comprehensive 25-item framework comprising five thematic clusters, items 11, 23, 24, and 25 are negatively formulated and must be reversed prior to computing composite scores.
    • Informed Data Retention / Self-Exclusion: In accordance with online research ethics protocols (Thielsch & Weltzin, 2012), respondents must confirm explicit consent at the conclusion of the survey; unconfirmed submissions are automatically filtered out.

Permissions & Fee and Test Year

The revised version of the Münster Questionnaire for the Evaluation of Lectures (MFE-Vr) was formally validated and published in its revised configuration between 2009 and 2012 by Meinald T. Thielsch and Gerrit Hirschfeld at the University of Münster (Department of Psychology). The instrument is considered an open-access, non-commercial public scientific assessment tool for academic, research, and non-profit higher education quality assurance purposes. No licensing fees or royalties are required for academic deployment. Educational institutions and independent researchers seeking to implement the German scale items or integrate the instrument into digital course evaluation systems may contact Dr. Meinald T. Thielsch ([email protected]) or Dr. Gerrit Hirschfeld ([email protected]) for benchmarking norms, digital templates, and institutional collaboration guidance.

References

The psychometric development, validation, and theoretical framework of the MFE-Vr are documented in the following scientific publications:

  • Bechler, C., & Thielsch, M. T. (2012). Evaluation im Bachelor- und Masterstudium: Gibt es einen Prüfungsstress-Effekt auf die studentische Lehrveranstaltungsbewertung? Zeitschrift für Evaluation, 11(2), 229–246.
  • Gediga, G., Hamborg, K.-C., & Willumeit, K. (2000). Das Kieler Evaluationsinstrument für Lehrveranstaltungen (KIEL). In K. D. Kubinger (Hrsg.), Moderne Methoden der Psychologischen Diagnostik (S. 145–162). Hogrefe.
  • Gollwitzer, M., & Schlotz, W. (2003). Das Trierer Inventar zur Lehrevaluation (TRIL). Diagnostica, 49(1), 34–44. https://doi.org/10.1026//0012-1924.49.1.34
  • Göritz, A. S., Soucek, R., & Bacher, J. (2005). Online-Befragungen in der Praxis: Erfahrungen und Empfehlungen. In A. S. Göritz (Hrsg.), Praxis der Online-Forschung (S. 11–28). Hogrefe.
  • Grabbe, C. (2003). Entwicklung und Erprobung eines Fragebogens zur Lehrevaluation [Unveröffentlichte Diplomarbeit]. Westfälische Wilhelms-Universität Münster.
  • Greenwald, A. G. (1997). Validity concerns and usefulness of student ratings of instruction. American Psychologist, 52(11), 1182–1186. https://doi.org/10.1037/0003-066X.52.11.1182
  • Haaser, M. (2006). Lehrevaluation im Fach Psychologie: Reanalyse und Weiterentwicklung der Münsteraner Evaluationsskalen [Unveröffentlichte Diplomarbeit]. Westfälische Wilhelms-Universität Münster.
  • Haaser, M., Thielsch, M. T., & Moeck, T. (2007). Webbasierte Lehrevaluation: Das Münsteraner System zur Online-Lehrevaluation (MOL). Zeitschrift für Medienpsychologie, 19(4), 140–151.
  • Hirschfeld, G., & Thielsch, M. T. (2009). Konfirmatorische Prüfung und Kürzung der Münsteraner Evaluationsskalen für Vorlesungen [Forschungsbericht]. Institut für Psychologie, Westfälische Wilhelms-Universität Münster.
  • Hu, L. t., & Bentler, P. M. (1999). Cutoff criteria for fit indexes in covariance structure analysis: Conventional criteria versus new alternatives. Structural Equation Modeling: A Multidisciplinary Journal, 6(1), 1–55. https://doi.org/10.1080/10705519909540118
  • Marsh, H. W. (1984). Students' evaluations of university teaching: Dimensionality, reliability, validity, potential biases and utility. Journal of Educational Psychology, 76(5), 707–754. https://doi.org/10.1037/0022-0663.76.5.707
  • Marsh, H. W. (2007). Students' evaluations of university teaching: A multidimensional perspective. In R. P. Perry & J. C. Smart (Eds.), The scholarship of teaching and learning in higher education: An evidence-based perspective (pp. 319–383). Springer. https://doi.org/10.1007/1-4020-5742-3_10
  • Mayer, R. E. (2002). Multimedia learning. Psychology of Learning and Motivation, 41, 85–139. https://doi.org/10.1016/S0079-7421(02)80005-6
  • Mutz, R. (2003). Wie valide sind studentische Lehrveranstaltungsbeurteilungen? Methodische Ansätze zur Bestimmung der Kriteriumsvalidität. Empirische Pädagogik, 17(3), 362–386.
  • Paas, F., Renkl, A., & Sweller, J. (2003). Cognitive load theory and instructional design: Recent developments. Educational Psychologist, 38(1), 1–4. https://doi.org/10.1207/S15326985EP3801_1
  • Rindermann, H. (1996). Untersuchungen zur Validität von Studentenurteilen: Eine Analyse des Heidelberger Inventars zur Lehrveranstaltungs-Evaluation (HILVE). Roderer.
  • Rindermann, H. (2003). Lehrevaluation an Hochschulen: Theoretische Grundlagen und empirische Befunde. Zeitschrift für Evaluation, 2(2), 193–220.
  • Rindermann, H. (2009). HILVE-II: Heidelberger Inventar zur Lehrveranstaltungs-Evaluation (2. Auflage). Hogrefe.
  • Schmidt, B., & Loßnitzer, T. (2010). Hochschuldidaktische Evaluation: Stand der Forschung und Überblick über deutschsprachige Instrumente. Zeitschrift für Hochschulentwicklung, 5(2), 1–21. https://doi.org/10.3217/zfhe-5-02/01
  • Souvignier, E., & Gold, A. (2002). Lehrevaluation: Feedback, Steuerung oder Forschung? Zeitschrift für Pädagogische Psychologie, 16(3/4), 169–180. https://doi.org/10.1024//1010-0652.16.3.169
  • Staufenbiel, T. (2000). Fragebogen zur Evaluation von Vorlesungen (FEVOR). Diagnostica, 46(4), 169–181. https://doi.org/10.1026//0012-1924.46.4.169
  • Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257–285. https://doi.org/10.1207/s15516709cog1202_4
  • Thielsch, M. T., Hirschfeld, G., & Moeck, T. (2010). Metaevaluation des Münsteraner Systems zur Online-Lehrevaluation [Forschungsbericht]. Psychologisches Institut, Westfälische Wilhelms-Universität Münster.
  • Thielsch, M. T., & Weltzin, M. (2012). Freiwilliger Selbstausschluss in Online-Befragungen: Auswirkungen auf die Datenqualität bei Lehrevaluationen. Zeitschrift für Evaluation, 11(1), 77–94.

Items of the Scale

Below are the authentic scale items in their original language as published in the standard psychometric validation studies, without modification or translation to preserve instrument validity and reliability:

Antwortvorgaben (Authentic Response Scale):

Für Items 8-19 7-stufiges Antwortformat mit den Optionen: 1 = "stimme gar nicht zu", 2 = "stimme nicht zu", 3 = "stimme eher nicht zu", 4 = "neutral", 5 = "stimme eher zu", 6 = "stimme zu" und 7 = "stimme vollkommen zu". Zusätzlich steht die Antwortoption "nicht sinnvoll beantwortbar" zur Verfügung.

Reverse Scoring Rules:

The instrument comprises five dimensions of 5 items each: 1. Didaktik (Didactics, items 1-5), 2. Anregung (Stimulation/Inspiration, items 6-10), 3. Interaktion (Interaction, items 11-15), 4. Struktur (Structure/Organization, items 16-20), 5. Rahmenbedingungen (Framework conditions, items 21-25). Items 11, 23, 24, and 25 are negatively formulated and reverse-coded.

  1. Die Vorlesungsinhalte werden durch Beispiele veranschaulicht.
  2. Die/der Dozent/in erläutert den praktischen Nutzen der behandelten Inhalte.
  3. Die/der Dozent/in setzt Medien (z.B. Beamer, Folien, Tafel) didaktisch sinnvoll ein.
  4. Die/der Dozent/in kann komplizierte Sachverhalte verständlich erklären.
  5. Wesentliche Inhalte werden durch Zusammenfassungen hervorgehoben.
  6. Die Vorlesung motiviert mich zur weiteren Auseinandersetzung mit dem Thema.
  7. Die/der Dozent/in versteht es, das Interesse für das Fachgebiet zu wecken.
  8. Die/der Dozent/in vermittelt die Inhalte auf engagierte und lebendige Weise.
  9. Die Vorlesung regt zum selbstständigen Nachdenken an.
  10. Die Vorlesung fördert mein Verständnis für Zusammenhänge im Fachgebiet.
  11. Die/der Dozent/in geht zu wenig auf Fragen und Beiträge der Studierenden ein. (-)
  12. Die/der Dozent/in nimmt die Studierenden und ihre Anliegen ernst.
  13. Es herrscht eine konstruktive und offene Arbeitsatmosphäre.
  14. Die/der Dozent/in zeigt Bereitschaft zur Unterstützung außerhalb der Vorlesung.
  15. Fragen und Diskussionsbeiträge von Studierenden sind ausdrücklich erwünscht.
  16. Der Aufbau der Vorlesung ist klar und nachvollziehbar gegliedert.
  17. Die Lernziele der Veranstaltung werden zu Beginn deutlich gemacht.
  18. Die einzelnen Vorlesungstermine bauen logisch aufeinander auf.
  19. Der zeitliche Rahmen der Vorlesung wird gut eingehalten.
  20. Die Relevanz der einzelnen Themen für das Gesamtziel wird deutlich.
  21. Die Raumgröße und Akustik sind für die Vorlesung angemessen.
  22. Die bereitgestellten Begleitmaterialien (z.B. Skripte, Folien) sind hilfreich.
  23. Der Schwierigkeitsgrad der Vorlesung ist unangemessen hoch. (-)
  24. Das Tempo der Vorlesung ist zu schnell. (-)
  25. Die Vorlesung ist durch störende Rahmenbedingungen (z.B. Lärm, Technikprobleme) beeinträchtigt. (-)
★

Rate This Scale

5.0 / 5 • 1 vote

Cite This Article

memjavad (2026, September 30). Münster Questionnaire for the Evaluation of Lectures – Revised (MFE-Vr). PSYCHOLOGICAL DATABASE. https://en.arabpsychology.com/scales/muenster-questionnaire-evaluation-lectures-revised-mfe-vr/
memjavad. “Münster Questionnaire for the Evaluation of Lectures – Revised (MFE-Vr).” PSYCHOLOGICAL DATABASE, 30 September 2026, https://en.arabpsychology.com/scales/muenster-questionnaire-evaluation-lectures-revised-mfe-vr/.
memjavad. “Münster Questionnaire for the Evaluation of Lectures – Revised (MFE-Vr).” PSYCHOLOGICAL DATABASE. September 30, 2026. https://en.arabpsychology.com/scales/muenster-questionnaire-evaluation-lectures-revised-mfe-vr/.