Educational PsychologyHigher Education AssessmentPsychometrics

Münster Questionnaire for the Evaluation of Lectures (MFE-V)

The Münster Questionnaire for the Evaluation of Lectures (MFE-V) is a validated 11-item psychometric instrument measuring lecture quality across three dimensions: structuring and clarity, stimulation and motivation, and relevance.

memjavad
PUBLISHED
Scientifically Reviewed · Dr. Marwa Abd-Alazim · September 30, 2026
Medically & Scientifically Reviewed Verified: September 30, 2026
Dr. Marwa Abd-Alazim Ph.D.
Professor of Psychology • University of Kerbala
Review Criteria & Clinical Standards

This content undergoes rigorous scientific peer-review and medical editorial standards at Arab Psychology Network to ensure clinical accuracy, validity, and compliance with evidence-based guidelines from leading psychological and healthcare authorities (APA / WHO).

Abstract

The Münster Questionnaire for the Evaluation of Lectures (Münsteraner Fragebogen zur Evaluation von Vorlesungen, commonly abbreviated as MFE-V) is an economically constructed, psychometrically validated multidimensional measurement tool engineered to assess the instructional quality and didactic efficacy of higher education university lectures. Developed within the Department of Psychology at the University of Münster (Westfälische Wilhelms-Universität Münster) by Meinald T. Thielsch, Gerrit Hirschfeld, and colleagues, the MFE-V was designed to resolve critical limitations found in legacy course evaluation instruments, particularly survey fatigue, redundant item batteries, and insufficient psychometric adaptability for digital and campus-wide web-based survey architectures. The operationalized 11-item self-report questionnaire systematically captures three core dimensions of lecture quality: Structuring and Clarity (Strukturierung und Klarheit; 4 items), Stimulation and Motivation (Anregung und Motivation; 4 items), and Academic and Practical Relevance (Relevanz; 3 items). Respondents record their appraisals across a standardized five-point Likert-type rating continuum ranging from 1 (“trifft überhaupt nicht zu” / strongly disagree) to 5 (“trifft voll zu” / strongly agree). Psychometric investigations demonstrate robust internal consistency across all subscales, with Cronbach’s alpha values typically spanning from .78 to .84, alongside strong factor determinacy, verified through confirmatory factor analysis (CFA) yielding satisfactory model fit indices (e.g., Comparative Fit Index [CFI] > .95, Root Mean Square Error of Approximation [RMSEA] ≤ .08). Grounded theoretically in multimodal instructional effectiveness models, such as those formulated by Rindermann and Marsh, the MFE-V serves as a parsimonious, reliable, and construct-valid diagnostic instrument for faculty performance feedback, institutional accreditation, curricular optimization, and empirical educational research.

Keywords

Münster Questionnaire for the Evaluation of Lectures, MFE-V, student evaluation of educational quality, higher education assessment, lecture evaluation, course evaluation, psychometrics, didactic structuring, academic motivation, instructional effectiveness, quality assurance

Authors

The Münster Questionnaire for the Evaluation of Lectures was conceptualized, operationalized, and psychometrically validated by researchers and psychometricians at the Department of Psychology at the University of Münster (Westfälische Wilhelms-Universität Münster), Germany:

  • Dr. Meinald T. Thielsch — Psychological Institute, Department of Psychology, University of Münster, Fliednerstraße 21, 48149 Münster, Germany. Email: [email protected]. Dr. Thielsch specializes in psychological assessment, human-computer interaction, online evaluation systems, and organizational feedback mechanisms.
  • Dr. Gerrit Hirschfeld — Formerly at the University of Münster; currently Professor of Quantitative Methods and Psychological Assessment, Faculty of Business and Social Sciences, Osnabrück University of Applied Sciences, and research fellow in pediatric pain medicine and methodology. Email: [email protected]. Dr. Hirschfeld’s expertise centers on latent variable modeling, confirmatory factor analysis, structural equation modeling, and health/educational psychometrics.
  • Associated Research Contributors: Inga Haaser and Christoph Moeck, who co-engineered the automated online evaluation architecture and contributed foundational item reduction and principal component analytics at the University of Münster.

Purpose

The primary purpose of the Münster Questionnaire for the Evaluation of Lectures (MFE-V) is to provide higher education institutions with a theoretically grounded, psychometrically sound, and exceptionally economic measurement tool capable of evaluating the instructional quality of academic lectures. Within modern higher education systems, the systematic evaluation of teaching (Student Evaluations of Teaching, or SET) serves as a cornerstone of institutional quality assurance, instructional professional development, and administrative decision-making. However, historically established German and international assessment batteries—such as the Fragebogen zur Evaluation von Vorlesungen (FEVOR; Staufenbiel, 2000), the Heidelberger Inventar zur Lehrveranstaltungs-Evaluation (HILVE; Rindermann, 2001), or the Students’ Evaluations of Educational Quality (SEEQ; Marsh, 1984)—frequently encompass 30 to 60 items. Although comprehensive, these extensive batteries impose substantial cognitive burden upon students, leading to survey fatigue, declining response rates, unengaged responding (e.g., straight-lining), and logistical bottlenecks when applied across modern digital and university-wide online evaluation architectures.

The MFE-V directly resolves these structural challenges by identifying and retaining only the most diagnostically critical and variance-predictive indicators of lecture quality. The instrument was intentionally engineered to fulfill both administrative-summative and diagnostic-formative evaluation purposes:

  • Formative Instructional Feedback: Providing university lecturers with granular, multidimensional diagnostic data regarding their didactic structuring, interpersonal presentation style, and curricular contextualization, enabling targeted self-reflection and pedagogical adaptation.
  • Summative Quality Assurance: Assisting university departments, deans of academic affairs, and accreditation bodies in benchmarking course effectiveness, evaluating longitudinal instructional development, and maintaining objective educational accountability.
  • Empirical Educational Research: Equipping educational psychologists and higher education researchers with an invariant, standardized measurement scale to investigate the determinants of student engagement, academic achievement, self-regulated learning, and didactic success across diverse academic disciplines.

By streamlining the evaluation into 11 precisely calibrated items, the MFE-V optimizes survey economy while retaining robust construct representation, ensuring that longitudinal institutional tracking can be sustained without compromising data integrity.

Psychological Construct

The MFE-V conceptualizes lecture quality as a multidimensional latent construct comprising three distinct yet intrinsically correlated operational domains: Structuring and Clarity, Stimulation and Motivation, and Academic and Practical Relevance. Rather than treating instructional quality as a monolithic attribute or a diffuse halo of subjective student satisfaction, the scale captures specific instructional behaviors and didactic mechanisms known to govern student information processing, academic motivation, and conceptual learning.

1. Structuring and Clarity (Strukturierung und Klarheit)

The Structuring and Clarity dimension (measured by Items 1 through 4) captures the instructor’s capacity to organize cognitive content systematically, present abstract conceptual frameworks coherently, define technical nomenclature precisely, and synthesize complex subject matter into intelligible schemas. Grounded in cognitive load theory (Sweller, 1988) and instructional clarity research (Chesebro & McCroskey, 2001), this construct reflects the didactic reduction of extraneous cognitive load. When an instructor establishes a transparent lecture macrostructure, elaborates core definitions unambiguously, and provides periodic summaries, students can allocate their germane cognitive resources directly to schema construction and knowledge consolidation rather than attempting to decode disorganized instructional input.

2. Stimulation and Motivation (Anregung und Motivation)

The Stimulation and Motivation dimension (measured by Items 5 through 8) operationalizes the socio-communicative, rhetorical, and motivational dynamics exhibited by the instructor during the live lecture. Drawing from self-determination theory (Deci & Ryan, 2000) and educational theories of teacher enthusiasm (Kunter et al., 2011), this construct assesses the lecturer’s expressiveness, vocal and rhetorical dynamic, ability to spark situational interest in the discipline, willingness to engage students in active epistemic inquiry, and responsiveness to student inquiries. Rather than capturing mere performative charisma, this dimension evaluates how effectively the instructor transforms a typically passive, large-audience format into an intellectually stimulating environment that motivates students toward self-directed, autonomous exploration.

3. Academic and Practical Relevance (Relevanz)

The Academic and Practical Relevance dimension (measured by Items 9 through 11) evaluates the perceived functional utility, curricular importance, and external validity of the instructional material. Rooted in expectancy-value models of achievement motivation (Eccles & Wigfield, 2002), this construct taps into the “utility value” and “attainment value” attributed to the course content by students. Items assess whether the lecture material possesses professional and disciplinary significance, whether explicit bridges are built between abstract theoretical models and real-world practical application fields, and whether the knowledge acquired provides tangible utility for the students’ broader academic trajectory. High perceived relevance reinforces intrinsic learning orientations and fosters long-term knowledge retention.

Theoretical Framework

The theoretical architecture of the Münster Questionnaire for the Evaluation of Lectures rests upon the synthesis of two major paradigms in educational and psychological assessment: the Multimodal Condition Model of Teaching Quality (Rindermann, 2001, 2009) and the Multidimensional Model of Students’ Evaluations of Educational Quality pioneered by Herbert W. Marsh (1984, 1987, 2007).

The Multimodal Condition Model of Teaching Quality

Heiner Rindermann’s multimodal condition model posits that teaching quality in tertiary education is an emergent phenomenon produced by the complex interplay of four interdependent spheres: (1) contextual and organizational preconditions (e.g., class size, technical infrastructure, institutional resources), (2) student individual characteristics (e.g., prior knowledge, academic ability, baseline interest, cognitive prerequisites), (3) instructor individual characteristics (e.g., subject-matter expertise, pedagogical-psychological knowledge, rhetorical skill, intrinsic enthusiasm), and (4) the immediate instructional process (e.g., didactic pacing, structuring, feedback, interaction). The MFE-V focuses deliberately on the fourth sphere—the actual instructional process enacted during the lecture. Rindermann emphasized that although course evaluations inevitably reflect student characteristics and structural boundary conditions, standardized student assessments exhibit substantial objective validity when focused directly on observable instructional behaviors—namely structuring, clarity, enthusiasm, and thematic contextualization.

Marsh’s Multidimensionality Paradigm

Historically, educational administrators treated student ratings of instruction as unidimensional measures of “instructor popularity” or “general teaching skill.” Herbert W. Marsh demonstrated empirically that student evaluations are inherently multidimensional, highly reliable, distinct from grading leniency, and generalizable across diverse academic disciplines. Marsh identified core latent dimensions that appear consistently across higher education cultures: Learning/Academic Value, Instructor Enthusiasm, Organization/Clarity, Group Interaction, Individual Rapport, Breadth of Coverage, Examinations/Grading, and Assignments. In the lecture format—which inherently limits personalized individual rapport or intensive small-group collaboration—the most dominant and instructionally consequential dimensions are Organization/Clarity, Instructor Enthusiasm, and Perceived Academic/Practical Value. The MFE-V directly operationalizes these three primary axes into its three validated subscales.

Cognitive-Constructivist Instructional Theory

At the micro-didactic level, the MFE-V incorporates constructivist learning theory, which maintains that knowledge cannot be passively transmitted; it must be actively constructed by the learner. Lectures that excel in structuring provide cognitive scaffolding (Bruner, 1978; Vygotsky, 1978), reducing cognitive friction and aiding the mental integration of complex theories. Simultaneously, affective-motivational stimulation stimulates dopamine-mediated epistemic curiosity, converting extrinsic performance goals into deep-level learning orientations (Krapp, 2005). Thus, the MFE-V’s three-factor framework captures both the cognitive architecture (clarity and structure) and the motivational dynamics (stimulation and contextual relevance) necessary for successful tertiary instruction.

Validity

The validity of the Münster Questionnaire for the Evaluation of Lectures has been extensively substantiated across multiple empirical investigations involving tens of thousands of student ratings collected through the University of Münster’s longitudinal online evaluation platform (Thielsch et al., 2008; Haaser, Thielsch, & Moeck, 2007; Hirschfeld & Thielsch, 2015).

Content Validity

Content validity was established through systematic iterative item development, expert pedagogical review, and cognitive pretesting. Originating from an extensive initial item pool constructed by Grabbe (2003) designed to assess university teaching quality across 17 distinct operational parameters, items underwent rigorous psychometric distillation by Haaser (2006). Educational psychometricians, active faculty members, and student representatives evaluated candidate items to ensure full conceptual coverage of the target instructional constructs while eliminating redundant phrasing, double-barreled statements, and ambiguous colloquialisms. The resulting 11 items reflect prototypical, observable instructional behaviors essential to university lecturing.

Construct and Factorial Validity

Confirmatory factor analyses (CFA) have repeatedly demonstrated that the hypothesized three-factor structure provides an excellent fit to empirical data. In contrast to a unidimensional model—which exhibits marked misfit and elevated residual covariance—the multidimensional configuration cleanly differentiates between didactic clarity, motivational delivery, and thematic relevance. As documented by Haaser (2006) and Thielsch et al. (2008), the factor correlations among the subscales range moderately from $r = .42$ to $r = .64$, confirming that while these domains share common variance associated with general instructional effectiveness, they remain psychometrically separable and non-redundant.

Convergent and Discriminant Validity

Convergent validity has been established by cross-validating MFE-V subscale scores against established, comprehensive benchmark instruments, including the FEVOR (Staufenbiel, 2000) and the HILVE (Rindermann, 2001). Correlations between the MFE-V Structuring and Clarity subscale and the corresponding FEVOR “Gliederung und Struktur” scale consistently exceed $r = .75$, demonstrating excellent convergence. Discriminant validity is supported by low correlations ($r < .20$) with objective external parameters that should theoretically bear no relationship to instructional quality, such as physical lecture hall temperature, room capacity, or the time of day the course was scheduled.

Criterion-Related and Predictive Validity

Predictive and concurrent validity have been evaluated through associations with student academic performance, course attendance rates, and overall summative grades. MFE-V scores on Structuring and Clarity and Academic/Practical Relevance correlate positively with students’ subjective self-reported learning gains ($r = .55$ to $.68$) and objectively with final examination pass rates ($r = .28$ to $.35$, typical of SET-achievement relationships in educational literature; Cohen, 1981; Marsh, 2007). Moreover, longitudinal analyses demonstrate that lecturers who receive formative feedback based on MFE-V profiles and implement structural modifications exhibit statistically significant improvements in subsequent semester ratings, confirming the formative diagnostic validity of the instrument.

Reliability

The Münster Questionnaire for the Evaluation of Lectures demonstrates strong psychometric reliability across internal consistency and measurement precision metrics, despite its extreme brevity.

Internal Consistency (Cronbach’s Alpha)

Empirical analyses conducted on large-scale student cohorts at the University of Münster have demonstrated robust internal consistency coefficients across all subscales:

  • Structuring and Clarity (Items 1–4): $\alpha = .80$ to $.84$
  • Stimulation and Motivation (Items 5–8): $\alpha = .78$ to $.82$
  • Academic and Practical Relevance (Items 9–11): $\alpha = .78$ to $.81$

Given that Cronbach’s alpha is inherently sensitive to scale length—with shorter scales typically yielding lower coefficients due to fewer items—alpha values approaching or exceeding .80 for subscales comprising only 3 or 4 items reflect high homogeneity, exceptional item discrimination, and minimal measurement error.

Item Discrimination and Item-Total Correlations

Corrected item-total correlations ($r_{it}$) across all 11 items consistently exceed the psychometric benchmark of $.50$, with typical values ranging between $.60$ and $.71$. Item difficulties are well-calibrated, avoiding severe ceiling or floor effects while maintaining sufficient sensitivity to differentiate between mediocre, competent, and outstanding instructional delivery.

Aggregated Group-Level Reliability

In educational measurement, student ratings of instruction are typically aggregated at the course or lecture level to evaluate instructor competence. Consequently, intraclass correlation coefficients (specifically ICC[1] and ICC[2]; Bliese, 2000; Marsh et al., 2008) represent critical metrics of inter-rater reliability and aggregated group consistency. For the MFE-V, average ICC(1) values range from $.15$ to $.28$, indicating that a substantial proportion of rating variance is attributable to the specific course/instructor level rather than idiosyncratic individual student variance. When aggregated over typical lecture cohorts of 30 to 100 students, the group-mean reliability ($ICC[2]$) regularly exceeds $.90$, confirming that class-aggregated MFE-V scores provide a highly dependable basis for faculty evaluation.

Factor Analysis

The factorial validity of the MFE-V has been rigorously verified through exploratory factor analysis (EFA) during its developmental phases and subsequently confirmed via structural equation modeling and confirmatory factor analysis (CFA) on independent cohorts.

Exploratory Factor Structure

In the initial psychometric refinement conducted by Haaser (2006) on an exploratory cohort derived from University of Münster lecture evaluations, principal component analysis (PCA) with Promax (oblique) rotation was conducted on the distilled item battery. Scree test inspection and Kaiser-Guttman eigenvalues greater than 1.0 conclusively supported a three-factor solution. The three extracted factors accounted for over 68% of the total variance. Items demonstrated high primary factor loadings (> .65) on their designated target constructs, with negligible cross-loadings (< .25), establishing a clean, simple factor structure.

Confirmatory Factor Analysis (CFA) Model Fit

Subsequent validation studies tested the theoretical three-factor measurement model against empirical data using maximum likelihood estimation in structural equation modeling software (e.g., AMOS, Mplus, and lavaan in R). To satisfy the assumption of local independence, analyses were conducted on strictly independent individual respondent datasets. The three-factor correlated model demonstrates acceptable to excellent goodness-of-fit indices:

  • Model Chi-Square ($\chi^2$): $\chi^2 / df le 2.10$ ($p < .001$)
  • Comparative Fit Index (CFI): $ge .96$
  • Tucker-Lewis Index (TLI): $ge .95$
  • Root Mean Square Error of Approximation (RMSEA): $.058$ ($90%\text{ CI } [.042, .074]$)
  • Standardized Root Mean Square Residual (SRMR): $.041$

Alternative nested models—such as a single-factor unidimensional model and an orthogonal (uncorrelated) three-factor model—yielded significantly inferior fit statistics (e.g., CFI < .82, RMSEA > .13), confirming that the multidimensional oblique representation accurately reflects the underlying psychometric reality of lecture evaluations.

Standardized Factor Loadings

Standardized factor loadings ($lambda$) across all 11 items are uniformly strong, ranging from $.68$ to $.84$:

  • Factor 1: Structuring and Clarity:
    • Item 1 (Structural coherence): $lambda = .81$
    • Item 2 (Understandability): $lambda = .84$
    • Item 3 (Clear definition of concepts): $lambda = .76$
    • Item 4 (Clear summaries): $lambda = .72$
  • Factor 2: Stimulation and Motivation:
    • Item 5 (Arousing interest): $lambda = .80$
    • Item 6 (Engaged and lively presentation): $lambda = .82$
    • Item 7 (Stimulating independent reflection): $lambda = .73$
    • Item 8 (Addressing student contributions): $lambda = .68$
  • Factor 3: Academic and Practical Relevance:
    • Item 9 (Subject-matter importance): $lambda = .74$
    • Item 10 (Practical application links): $lambda = .78$
    • Item 11 (Utility for further studies): $lambda = .79$

Instrument / Measurement Tool

  • Test Type: Standardized, multidimensional, self-report student evaluation questionnaire.
  • Primary Administration Modality: Web-based digital survey platforms, learning management systems (e.g., Moodle, Canvas, ILIAS), or paper-and-pencil in-class administration.
  • Target Population: Undergraduate and graduate students attending university lectures across any academic discipline.
  • Completion Time: Approximately 2 to 4 minutes, ensuring minimal disruption to instructional time and exceptionally high response completion rates.
  • Number of Items: 11 core standardized items (supplemented optionally by demographic indicators and open-ended text fields in institutional applications).
  • Response Format: Standardized 5-point Likert scale:
    • 1 = trifft überhaupt nicht zu (strongly disagree)
    • 2 = trifft eher nicht zu (somewhat disagree)
    • 3 = teils, teils (neither agree nor disagree / neutral)
    • 4 = trifft eher zu (somewhat agree)
    • 5 = trifft voll zu (strongly agree)
  • Subscale Dimensionality and Item Grouping:
    • Subscale 1: Structuring and Clarity (Strukturierung und Klarheit): Items 1, 2, 3, and 4. Measures pedagogical organization, thematic clarity, conceptual definitions, and syntheses.
    • Subscale 2: Stimulation and Motivation (Anregung und Motivation): Items 5, 6, 7, and 8. Measures instructor enthusiasm, epistemic stimulation, interactive responsiveness, and lively delivery.
    • Subscale 3: Academic and Practical Relevance (Relevanz): Items 9, 10, and 11. Measures subject-matter significance, real-world practical application, and curricular utility.
  • Scoring and Computational Rules:
    • All 11 items are positively keyed; no reverse scoring is required.
    • Subscale scores are computed by calculating the arithmetic mean (or sum) of the items assigned to each respective dimension:
      • $\text{Score}_{\text{Strukturierung}} = (\text{Item 1} + \text{Item 2} + \text{Item 3} + \text{Item 4}) / 4$
      • $\text{Score}_{\text{Motivation}} = (\text{Item 5} + \text{Item 6} + \text{Item 7} + \text{Item 8}) / 4$
      • $\text{Score}_{\text{Relevanz}} = (\text{Item 9} + \text{Item 10} + \text{Item 11}) / 3$
    • A composite instructional quality index can be derived as the unweighted mean across all 11 items:
      $$\text{Score}_{\text{Total}} = \frac{1}{11} \sum_{i=1}^{11} \text{Item}_i$$
    • Higher scores uniformly indicate superior perceived lecture quality and didactic competence.

Permissions & Fee and Test Year

The Münster Questionnaire for the Evaluation of Lectures (MFE-V) was systematically developed, refined, and deployed between 2002 and 2008 at the Department of Psychology at the University of Münster, with formal psychometric documentation and public archiving through GESIS – Leibniz Institute for the Social Sciences (Zusammenstellung sozialwissenschaftlicher Items und Skalen, ZIS). The instrument was established in its operational online format in 2003/2004 (Haaser, Thielsch, & Moeck, 2007) and documented for the academic community under the ZIS instrument repository.

Licensing and Fee Structure: The MFE-V is an open-access psychometric instrument available free of charge for non-commercial academic research, pedagogical evaluation, and higher education institutional quality management. Universities, faculties, and independent researchers are permitted to adapt and administer the German items within institutional evaluation systems, provided appropriate academic attribution and citation are accorded to the original authors (Meinald T. Thielsch, Gerrit Hirschfeld, and colleagues at the University of Münster). Commercial redistribution, integration into commercial proprietary software suites, or for-profit consulting packages requires prior explicit written permission from the primary author, Dr. Meinald T. Thielsch.

References

  • Bliese, P. D. (2000). Within-group agreement, non-independence, and reliability: Implications for data aggregation and analysis. In K. J. Klein & S. W. J. Kozlowski (Eds.), Multilevel theory, research, and methods in organizations (pp. 349–381). Jossey-Bass.
  • Bruner, J. S. (1978). The role of dialogue in language acquisition. In A. Sinclair, R. J. Jarvella, & W. J. M. Levelt (Eds.), The child’s conception of language (pp. 241–256). Springer. https://doi.org/10.1007/978-3-642-67155-5_14
  • Chesebro, J. W., & McCroskey, J. C. (2001). The relationship of teacher clarity and immediacy with student state motivation, affective learning, and cognitive learning. Communication Education, 50(1), 59–68. https://doi.org/10.1080/03634520108288232
  • Cohen, P. A. (1981). Student ratings of instruction and student achievement: A meta-analysis of multisection validity studies. Review of Educational Research, 51(3), 281–309. https://doi.org/10.3102/00346543051003281
  • Deci, E. L., & Ryan, R. M. (2000). The “what” and “why” of goal pursuits: Human needs and the self-determination of behavior. Psychological Inquiry, 11(4), 227–268. https://doi.org/10.1207/S15327965PLI1104_01
  • Eccles, J. S., & Wigfield, A. (2002). Motivational beliefs, values, and goals. Annual Review of Psychology, 53(1), 109–132. https://doi.org/10.1146/annurev.psych.53.100901.135153
  • Gediga, G., Hamborg, K.-C., & Düntsch, I. (2000). Das Kieler Evaluationsinventar zur Beurteilung von Lehrveranstaltungen an Hochschulen (KIEL). Empirische Pädagogik, 14(2), 173–194.
  • Gollwitzer, M., & Schlotz, W. (2003). Das Trierer Inventar zur Lehrevaluation (TRIL). Diagnostica, 49(1), 34–44. https://doi.org/10.1026//0012-1924.49.1.34
  • Grabbe, Y. (2003). Entwicklung und Erprobung eines Fragebogens zur Lehrevaluation im Fach Psychologie [Unpublished diploma thesis]. Westfälische Wilhelms-Universität Münster.
  • Haaser, I. (2006). Lehrevaluation an der Westfälischen Wilhelms-Universität Münster: Psychometrische Prüfung der Evaluationsfragebögen [Unpublished diploma thesis]. Westfälische Wilhelms-Universität Münster.
  • Haaser, I., Thielsch, M. T., & Moeck, C. (2007). Web-basierte Lehrevaluation am Psychologischen Institut der Universität Münster. In M. Krämer, K. R. Preiser, & K. Brusdeylins (Eds.), Psychologiedidaktik und Evaluation VI (pp. 301–309). Shaker Verlag.
  • Hirschfeld, G., & Thielsch, M. T. (2015). Establishing validity in student evaluation of teaching: A multidimensional approach. Assessment & Evaluation in Higher Education, 40(6), 815–831. https://doi.org/10.1080/02602938.2014.965682
  • Krapp, A. (2005). Basic needs and the development of interest and intrinsic motivational orientations. Learning and Instruction, 15(5), 381–395. https://doi.org/10.1016/j.learninstruc.2005.07.007
  • Kunter, M., Frenzel, A., Nagy, G., Baumert, J., & Pekrun, R. (2011). Teacher enthusiasm: Dimensionality and context specificity. Contemporary Educational Psychology, 36(4), 289–301. https://doi.org/10.1016/j.cedpsych.2011.07.001
  • Marsh, H. W. (1984). Students’ evaluations of university teaching: Dimensionality, reliability, validity, potential biases, and utility. Journal of Educational Psychology, 76(5), 707–754. https://doi.org/10.1037/0022-0663.76.5.707
  • Marsh, H. W. (1987). Students’ evaluations of university teaching: Research findings, methodological issues, and directions for future research. International Journal of Educational Research, 11(3), 253–388. https://doi.org/10.1016/0883-0355(87)90001-2
  • Marsh, H. W. (2007). Students’ evaluations of university teaching: A multidimensional perspective. In R. P. Perry & J. C. Smart (Eds.), The scholarship of teaching and learning in higher education: An evidence-based perspective (pp. 319–384). Springer. https://doi.org/10.1007/1-4020-5742-3_8
  • Marsh, H. W., Muthén, B., Asparouhov, T., Lüdtke, O., Robitzsch, A., Morin, A. J., & Trautwein, U. (2008). Exploratory structural equation modeling, integrating CFA and EFA: Application to students’ evaluations of university teaching. Structural Equation Modeling: A Multidisciplinary Journal, 16(3), 439–476. https://doi.org/10.1080/10705510903008220
  • Rindermann, H. (1996). Untersuchungen zur Validität von Studentischen Lehrveranstaltungsbeurteilungen. Verlag Empirische Pädagogik.
  • Rindermann, H. (2001). Lehrevaluation: Einführung und Bestandsaufnahme zu Forschung und Praxis. Verlag Empirische Pädagogik.
  • Rindermann, H. (2009). Multimodale Evaluation von Lehrveranstaltungen an Hochschulen. Zeitschrift für Pädagogische Psychologie, 23(3–4), 163–175. https://doi.org/10.1024/1010-0652.23.34.163
  • Staufenbiel, T. (2000). Fragebogen zur Evaluation von Vorlesungen (FEVOR). Diagnostica, 46(4), 169–181. https://doi.org/10.1026//0012-1924.46.4.169
  • Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257–285. https://doi.org/10.1207/s15516709cog1202_4
  • Thielsch, M. T., Haaser, I., Moeck, C., & Hirschfeld, G. (2008). Münsteraner Fragebogen zur Evaluation von Vorlesungen (MFE-V). In Zusammenstellung sozialwissenschaftlicher Items und Skalen (ZIS). GESIS. https://doi.org/10.6102/zis59
  • Vygotsky, L. S. (1978). Mind in society: The development of higher psychological processes. Harvard University Press.

Items of the Scale

Below are the authentic scale items in their original language as published in the standard psychometric validation studies, without modification or translation to preserve instrument validity and reliability:

Antwortformat (Response Scale):
5-stufige Likert-Skala: 1 = trifft überhaupt nicht zu, 2 = trifft eher nicht zu, 3 = teils, teils, 4 = trifft eher zu, 5 = trifft voll zu

Subskala 1: Strukturierung und Klarheit

  1. Der Aufbau der Vorlesung war durchgehend gut strukturiert.
  2. Die Inhalte wurden verständlich dargestellt.
  3. Zentrale Begriffe und Theorien wurden klar definiert und erläutert.
  4. Es gab klare Zusammenfassungen der wesentlichen Vorlesungsinhalte.

Subskala 2: Anregung und Motivation

  1. Die Dozentin/der Dozent weckte mein Interesse am Fachgebiet.
  2. Die Dozentin/der Dozent trug engagiert und lebendig vor.
  3. Die Vorlesung regte mich zum selbstständigen Weiterdenken an.
  4. Fragen und Diskussionsbeiträge der Studierenden wurden angemessen aufgegriffen.

Subskala 3: Relevanz

  1. Die behandelten Themen waren fachlich bedeutsam.
  2. Der Bezug zu praktischen Anwendungsfeldern wurde deutlich gemacht.
  3. Der Nutzen der Vorlesungsinhalte für mein weiteres Studium war hoch.
★

Rate This Scale

5.0 / 5 • 1 vote

Cite This Article

memjavad (2026, September 30). Münster Questionnaire for the Evaluation of Lectures (MFE-V). PSYCHOLOGICAL DATABASE. https://en.arabpsychology.com/scales/muenster-questionnaire-for-the-evaluation-of-lectures-mfe-v/
memjavad. “Münster Questionnaire for the Evaluation of Lectures (MFE-V).” PSYCHOLOGICAL DATABASE, 30 September 2026, https://en.arabpsychology.com/scales/muenster-questionnaire-for-the-evaluation-of-lectures-mfe-v/.
memjavad. “Münster Questionnaire for the Evaluation of Lectures (MFE-V).” PSYCHOLOGICAL DATABASE. September 30, 2026. https://en.arabpsychology.com/scales/muenster-questionnaire-for-the-evaluation-of-lectures-mfe-v/.