Abstract
The Additional Module Basic Texts – Münster Questionnaire for Evaluation (German: Münsteraner Fragebogen zur Evaluation – Zusatzmodul Basistexte, abbreviated as MFE-ZBa) is an economic, psychometrically validated modular survey instrument designed to evaluate the didactic integration, comprehensibility, and pedagogical utility of foundational reading texts in higher education courses. Developed at the University of Münster (Westfälische Wilhelms-Universität Münster) as part of the broader Münster Questionnaire for Evaluation (MFE) system, the MFE-ZBa addresses a critical gap in traditional course evaluation instruments, which frequently emphasize overall lecturer performance while overlooking specific didactic media. Comprising five items rated on a 7-point Likert scale (ranging from 1 = “stimme gar nicht zu” / strongly disagree to 7 = “stimme vollkommen zu” / strongly agree, alongside an explicit non-applicability opt-out option), the instrument assesses student perceptions regarding text clarity, alignment with course learning objectives, cognitive scaffolding, reasonable time expenditure, and overall perceived text quality. Psychometric evaluations conducted across university cohorts (N = 242) demonstrate solid internal consistency (Cronbach’s alpha = .79) and an excellent fit for a unidimensional measurement model via confirmatory factor analysis (RMSEA = .08, CFI = .98, TLI = .93). Construct validation confirms robust convergent associations with lecturer didactic competence (r = .63 in seminars; r = .52 in lectures), instructional materials (r = .62 in seminars; r = .45 in lectures), and subjective learning gains (r = .56 in seminars; r = .39 in lectures), alongside divergent validity demonstrated through significant negative correlations with student academic overburdening (r = -.53 in seminars; r = -.40 in lectures). Furthermore, the scale demonstrates discriminative sensitivity across disparate academic courses (F = 2.14, p < .01, η² = .12), establishing the MFE-ZBa as a rigorous diagnostic tool for quality assurance, instructional development, and educational research in tertiary education.
Keywords
teaching evaluation, course evaluation, higher education, student evaluation of teaching, MFE-ZBa, instructional reading, educational quality assurance, psychometrics, constructive alignment, academic reading
Authors
The Münster Questionnaire for Evaluation and its specific supplementary modules were conceptualized, validated, and implemented by researchers at the Department of Psychology, University of Münster (Westfälische Wilhelms-Universität Münster, Germany):
- Dr. Dipl.-Psych. Meinald T. Thielsch — Westfälische Wilhelms-Universität Münster, Institut für Psychologie, Fliednerstraße 21, 48149 Münster, Germany. E-Mail: [email protected]; Project Website: www.uni-muenster.de/PsyEval.
- B.Sc. Ina Stegemöller — Westfälische Wilhelms-Universität Münster, Institut für Psychologie, Fliednerstraße 21, 48149 Münster, Germany.
Purpose
Student Evaluation of Teaching (SET) serves as a cornerstone of institutional quality assurance, instructional development, and curriculum reform across European and international universities. As delineated by foundational research in tertiary pedagogy (Rindermann, 1996), systematic course ratings fulfill multiple strategic functions: they stimulate the reflective enhancement of pedagogical competency among instructors, pinpoint instructional strengths and deficits across departments, foster dialogue between faculty and students, inform resource allocation, and direct targeted academic faculty development. However, conventional course evaluation batteries typically rely on extensive, generalized questionnaires that assess global lecture parameters, thereby obscuring the nuanced efficacy of specific instructional components.
The primary purpose of the Additional Module Basic Texts (MFE-ZBa) is to provide an economically targeted, modular evaluation of foundational readings (Basistexte) assigned in university lectures and seminars. In contemporary university curricula—particularly following the modularized Bachelor and Master restructuring under the Bologna Process—student workload and self-directed preparation through mandatory literature constitute substantial proportions of total academic credit allocations (European Credit Transfer and Accumulation System, ECTS). Despite the central role of preparatory literature in determining student comprehension during in-class discussions, standardized SET inventories rarely isolate whether assigned texts are comprehensible, aligned with session goals, pedagogically scaffolded, or disproportionately burdensome.
The MFE-ZBa addresses these operational imperatives by serving both formative and summative evaluative functions:
- Formative Pedagogical Diagnostic: Instructors receive granular, actionable feedback concerning whether their selected preparatory texts clarify complex conceptual frameworks or introduce cognitive overload, enabling informed decisions to retain, replace, or scaffold specific readings in subsequent academic terms.
- Summative Quality Management: Academic departments and accreditation bodies obtain standardized metrics assessing whether reading loads comply with curricular standards without inducing excessive student strain.
- Workload and Stress Mitigation: Administered through web-based evaluation platforms at the conclusion of the semester—a period notorious for student assessment anxiety and time constraints—the modular structure allows instructors to deploy the 5-item scale selectively, eliminating irrelevant survey items and minimizing survey fatigue.
Psychological Construct
The latent construct assessed by the MFE-ZBa is the Perceived Instructional Quality and Didactic Efficacy of Assigned Course Literature (Lehrqualität von Basistexten). Within higher education psychology, assigned reading materials are not neutral textual artifacts; rather, they represent central instructional interventions designed to facilitate cognitive schema acquisition, prior knowledge activation, and active student engagement prior to or following formal lecture sessions. The construct encapsulates several interrelated pedagogical and cognitive facets operationalized through a coherent, unidimensional measurement model:
1. Textual Comprehensibility and Readability (Item 1)
Textual comprehensibility reflects the degree to which instructional literature matches the cognitive developmental stage, reading competence, and domain-specific vocabulary of the student cohort. In line with psycholinguistic models of text comprehension, assigned literature must balance conceptual sophistication with structural clarity, coherent thematic progression, and transparent argumentation. When texts exhibit excessive syntactic opacity or unexplained disciplinary jargon, students experience high extrinsic cognitive load, hindering deep comprehension.
2. Curricular Alignment and Goal Transparency (Item 2)
Curricular alignment denotes the explicit, transparent link between assigned preparatory texts and the formal learning objectives (Ziele der Veranstaltung) established for classroom sessions. Drawing upon educational psychology principles of purposeful learning, students require clear structural orientation demonstrating how individual readings contribute directly to examination benchmarks, case discussions, or overarching disciplinary proficiencies.
3. Cognitive Utility and Learning Facilitation (Item 3)
This facet assesses the subjective utility and explanatory power of the readings in consolidating instructional subject matter. High-quality foundational literature operates as an effective cognitive scaffold, bridging abstract theoretical concepts presented by the lecturer and the student’s internal cognitive schemata. When texts are well-chosen, students perceive reading tasks not as perfunctory homework, but as an indispensable facilitator of content mastery.
4. Cognitive Workload and Time Investment Proportionality (Item 4, Reverse-Coded)
Academic reading inevitably demands temporal and cognitive resources. This dimension captures the perceived reasonableness of the invested time relative to pedagogical return. Disproportionate or punitive time requirements trigger student disengagement, surface-level skimming strategies, and heightened academic stress. Within the MFE-ZBa framework, this facet is framed as an inverted indicator of instructional design quality, highlighting whether assigned volumes exceed realistic credit-hour allocations.
5. Global Evaluative Satisfaction (Item 5)
Representing an overarching summative appraisal, this dimension captures the student’s consolidated evaluation of the literature’s overall didactic caliber. It reflects the synthesis of utility, accessibility, engagement, and aesthetic-didactic presentation.
Theoretical Framework
The architectural foundation of the MFE-ZBa is rooted in three convergent theoretical frameworks within educational and cognitive psychology:
1. Cognitive Load Theory (CLT)
Developed by John Sweller and colleagues, Cognitive Load Theory posits that human working memory capacity is strictly limited. Instructional media impose three distinct forms of cognitive load:
- Intrinsic cognitive load: Demanded by the inherent conceptual complexity of the material itself.
- Extraneous cognitive load: Generated by poor instructional presentation, convoluted textual phrasing, or disjointed organizational structure.
- Germane cognitive load: Devoted to genuine schema construction, organization, and long-term memory integration.
Within this framework, foundational texts in tertiary courses must be curated to minimize extraneous cognitive load (high comprehensibility, explicit objectives) while providing sufficient scaffolding to channel student cognitive effort into germane learning processes. The MFE-ZBa directly measures the manifestation of these cognitive loads: Item 1 and Item 3 assess germane schema support, whereas Item 4 identifies elevated extraneous burden caused by poorly curated literature.
2. The Theory of Constructive Alignment
Formulated by John B. Biggs, the Constructive Alignment paradigm asserts that optimal learning environments occur when intended learning outcomes, teaching and learning activities, and assessment tasks are fully congruent. In higher education seminars and flipped-classroom lectures, reading basic texts represents the primary pre-class learning activity. If an incongruity exists between assigned literature and in-class discussions, students experience disorientation, leading to fragmented learning approaches. Item 2 of the MFE-ZBa explicitly captures this systemic alignment.
3. Multidimensional Models of Student Evaluation of Teaching (SET)
Pioneering SET theorists, notably Herbert W. Marsh (1984), demonstrated that student evaluations of instructional quality are inherently multidimensional, stable, and functionally valid reflections of classroom dynamics. Marsh emphasized that teaching effectiveness cannot be distilled into a single, homogeneous score without losing diagnostic utility. Concurrently, Heiner Rindermann (1996) emphasized that diagnostic evaluations must capture discrete instructional parameters, including didactic competence, interaction quality, pacing, and learning resources.
Recognizing the limitations of monolithic 30- to 50-item evaluation surveys, the developers at the University of Münster (Grabbe, 2003; Moeck & Thielsch, 2004) conceptualized a modular SET architecture. Rather than imposing invariant reading-material questions upon courses devoid of literature assignments, the core questionnaires (MFE-Sr for seminars, MFE-Vr for lectures) maintain high parsimony, while modular add-ons like the MFE-ZBa are activated conditionally based on instructional design. This preserves student survey motivation and satisfies the psychometric requirements of modern web-based survey methodologies (Thielsch & Weltzin, 2012).
Validity
Empirical validation of the MFE-ZBa was conducted using cross-sectional datasets gathered across multiple academic semesters at the University of Münster. As noted by Marsh (1984) and Rindermann (1996), establishing construct validity for specialized SET modules requires examining whether the targeted sub-dimension demonstrates theoretically predictable convergence with broader instructional indicators while maintaining discriminant uniqueness.
Convergent Validity
Scale validity was investigated by correlating the aggregate MFE-ZBa scale scores with established subscales of the standardized Münster core evaluation inventories: the Seminar Module (MFE-Sr) and the Lecture Module (MFE-Vr). Pearson correlation analyses revealed substantial, statistically significant positive relationships with central pedagogical dimensions:
- Lecturer Competence & Didactics (Dozent & Didaktik): Substantial positive correlations emerged in both seminar formats (r = .63, p < .01) and lecture formats (r = .52, p < .01). These coefficients affirm that students perceive the skillful curation and didactic contextualization of literature as an integral expression of the instructor’s general pedagogical expertise.
- Course Materials (Materialien): Moderate to high correlations were observed with the core instructional materials scale (r = .62 in seminars; r = .45 in lectures), demonstrating robust convergent overlap while confirming that basic texts represent a distinct subset of broader physical and digital media.
- Subjective Learning Gain (Selbst eingeschätzter Lernerfolg): Basic text evaluation correlated significantly with student-rated academic attainment (r = .56 in seminars; r = .39 in lectures), substantiating the premise that didactic literature directly supports cognitive mastery.
- Overall Course Evaluation (Gesamtbeurteilung): Global course satisfaction was moderately to strongly associated with basic text ratings (r = .56 in seminars; r = .35 in lectures).
- Course Recommendation (Weiterempfehlung): In seminar environments, text ratings correlated positively with students’ willingness to recommend the course to peers (r = .29, p < .01).
- Peer Interaction Quality (Teilnehmer): In interactive seminar settings, basic text ratings correlated moderately with peer interaction dynamics (r = .33, p < .01), illustrating that shared foundational reading fosters more collaborative in-class discussions.
Discriminant and Divergent Validity
Divergent validity was substantiated through associations with the MFE core scale measuring student cognitive and temporal overburdening (Überforderung):
- A substantial negative correlation was observed in seminars (r = -.53, p < .01).
- A moderate-to-high negative correlation was confirmed in lecture formats (r = -.40, p < .01).
These divergent coefficients demonstrate that well-chosen, comprehensible literature actively alleviates subjective student distress and disorientation, disconfirming any hypothesis that high reading quality is conflated with unmanageable academic pressure.
Discriminative Criterion Validity (Between-Course Variance)
A critical psychometric criterion for any course evaluation tool is its capacity to differentiate reliably between distinct instructional settings rather than merely capturing individual student response styles. A one-way analysis of variance (ANOVA) conducted across diverse courses using course identifier as the independent factor and aggregate MFE-ZBa scores as the dependent variable revealed significant between-course differentiation:
F(15, 226) = 2.14, p < .01, η² = .12
An effect size of η² = .12 represents a medium-to-large instructional effect, confirming that the MFE-ZBa is sensitive to variations in literature curation across instructors and departmental curricula.
Reliability
Reliability analyses conducted on the validation sample (N = 242) demonstrate satisfactory psychometric internal consistency for an applied educational diagnostic tool:
- Internal Consistency: The scale yields an overall Cronbach’s alpha of α = .79. In applied educational research and group-level course evaluations—where aggregated class means represent the primary unit of analysis—coefficients approaching or exceeding .80 satisfy established psychometric standards (Nunnally & Bernstein, 1994; Marsh, 1984).
- Corrected Item-Total Correlations (Trennschärfen, rit): All five items demonstrate acceptable to strong item-total discrimination, ranging from .45 to .64:
| Item | Mean (M) | SD | Item-Total r (T) | Factor Loading (λ) | α if Item Deleted |
|---|---|---|---|---|---|
| Item 1 (Verständlichkeit) | 5.41 | 1.21 | .52 | .62 | .77 |
| Item 2 (Bezug zu Zielen) | 5.83 | 1.35 | .64 | .75 | .73 |
| Item 3 (Verständnishilfe) | 5.30 | 1.59 | .45 | .52 | .79 |
| Item 4 (Zeitaufwand; recoded) | 5.04 | 1.50 | .62 | .70 | .73 |
| Item 5 (Gesamtbewertung) | 5.52 | 1.50 | .64 | .77 | .73 |
Decomposition of α-if-item-deleted statistics indicates that no item removal increases internal consistency beyond the overall .79 benchmark, corroborating that all five indicators contribute proportionally to true score variance.
Factor Analysis
The structural dimensionality of the MFE-ZBa was examined using Confirmatory Factor Analysis (CFA) utilizing maximum likelihood estimation in AMOS, based upon initial exploratory analyses from antecedent semesters (Winter Semester 2008/09 and 2009/10). The hypothesized measurement model posited a single, unidimensional latent factor representing the perceived didactic quality of basic literature.
Model Fit Parameters
The unidimensional model demonstrated acceptable-to-excellent empirical fit across conventional global fit criteria:
- Chi-Square (χ²): χ² = 11.51, degrees of freedom (df) = 5, χ²/df = 2.30. The relative chi-square ratio falling below 3.0 indicates an adequate data-to-model fit.
- Comparative Fit Index (CFI): .98 (exceeding the conservative ≥ .95 threshold for good fit).
- Tucker-Lewis Index (TLI): .93 (exceeding standard acceptable fit thresholds of ≥ .90).
- Root Mean Square Error of Approximation (RMSEA): .08 (reflecting reasonable error of approximation according to Browne & Cudeck standards).
Standardized Factor Loadings
All standardized factor loadings (λ) loaded substantially and statistically significantly (p < .001) onto the common latent factor:
- Item 1: λ = .62
- Item 2: λ = .75
- Item 3: λ = .52
- Item 4 (recoded): λ = .70
- Item 5: λ = .77
The strong loadings of Item 2 (.75), Item 4 (.70), and Item 5 (.77) emphasize that student evaluations of course texts are anchored in clear curricular relevance, reasonable workload requirements, and overall pedagogical utility.
Instrument / Measurement Tool
The MFE-ZBa is structured as follows:
- Tool Name: Münster Questionnaire for Evaluation – Basic Texts Module (MFE-ZBa)
- Original Language: German (Münsteraner Fragebogen zur Evaluation – Zusatzmodul Basistexte)
- Instrument Type: Standardized self-report rating scale / modular institutional survey
- Number of Items: 5 items
- Response Format: Fully labeled 7-point Likert rating scale:
- 1 = “stimme gar nicht zu” (Strongly disagree)
- 2 = “stimme nicht zu” (Disagree)
- 3 = “stimme eher nicht zu” (Somewhat disagree)
- 4 = “neutral” (Neutral)
- 5 = “stimme eher zu” (Somewhat agree)
- 6 = “stimme zu” (Agree)
- 7 = “stimme vollkommen zu” (Strongly agree)
- Zusatzoption: “nicht sinnvoll beantwortbar” (Cannot be evaluated / not applicable; treated as missing data)
- Target Population: Undergraduate and graduate university students in tertiary education courses (lectures, seminars, colloquia).
- Scoring Instructions:
- Reverse Scoring: Item 4 is negatively phrased (“Der Zeitaufwand für die Bearbeitung der Basistexte war unangemessen hoch”) and must be reverse-coded prior to scale aggregation (recoded score = 8 – raw score; i.e., 1 → 7, 2 → 6, 3 → 5, 4 → 4, 5 → 3, 6 → 2, 7 → 1).
- Aggregation: Given confirmed unidimensionality, an overall Basic Texts Quality Score is derived by computing the arithmetic mean across all five items (or summing across items, theoretical range: 5 to 35). Higher composite scores signify superior didactic curation, comprehensibility, and pedagogical alignment.
- Missing Values: Opt-out selections (nicht sinnvoll beantwortbar) are excluded pairwise from individual scale score computations.
Permissions & Fee and Test Year
The Münster Questionnaire for Evaluation (MFE) system and its modular components were systematically constructed and evaluated beginning in the 2000/2001 academic cycle, with online platform deployment initiated in 2003/2004 (Haaser, Thielsch, & Moeck, 2007). Comprehensive psychometric revisions and standardization of the additional modules, including the MFE-ZBa, were completed in 2010.
The instrument is made available for academic research, university teaching evaluation, and institutional quality assurance free of charge (Open Access for non-commercial academic purposes). Institutional usage, automated software integrations, and inquiries regarding adaptation permissions should be addressed to the primary developer:
Dr. Dipl.-Psych. Meinald T. Thielsch
Westfälische Wilhelms-Universität Münster, Institut für Psychologie
Fliednerstraße 21, 48149 Münster, Germany
E-Mail: [email protected]
Portal: www.uni-muenster.de/PsyEval
References
- Biggs, J. (1996). Enhancing teaching through constructive alignment. Higher Education, 32(3), 347–364. https://doi.org/10.1007/BF00138871
- Göritz, A. S., Soucek, R., & Bacher, J. (2005). Internet-based student evaluation of teaching: A review and empirical comparison. Zeitschrift für Medienpsychologie, 17(3), 88–99. https://doi.org/10.1026/1617-6383.17.3.88
- Grabbe, Y. (2003). Entwicklung und Erprobung eines Fragebogens zur Seminar-Evaluation [Unpublished diploma thesis]. Westfälische Wilhelms-Universität Münster, Münster, Germany.
- Haaser, A., Thielsch, M. T., & Moeck, A. (2007). Web-basierte Lehrevaluation: Konzeption, Umsetzung und Praxiserfahrungen an der Universität Münster. In M. Krämer, S. Preiser, & K. Brusdeylins (Eds.), Psychologiedidaktik und Evaluation VI (pp. 317–326). Shaker.
- Marsh, H. W. (1984). Students’ evaluations of university teaching: Dimensionality, reliability, validity, potential biases, and utility. Journal of Educational Psychology, 76(5), 707–754. https://doi.org/10.1037/0022-0663.76.5.707
- Moeck, A., & Thielsch, M. T. (2004). Münsteraner Fragebogen zur Evaluation (MFE) – Handanweisung für Zusatzmodule. Psychologisches Institut IV, Universität Münster.
- Rindermann, H. (1996). Untersuchungen zur Validität von Studentischen Lehrevaluationen. Empirische Pädagogik.
- Schmidt, F. T., & Loßnitzer, T. (2010). Instrumente zur Erfassung der Lehrqualität an Hochschulen: Eine Übersicht. Universität Heidelberg.
- Sweller, J. (2010). Element interactivity and intrinsic, extraneous, and germane cognitive load. Educational Psychology Review, 22(2), 123–138. https://doi.org/10.1007/s10648-010-9128-5
- Thielsch, M. T., & Weltzin, S. (2012). Online-Lehrevaluation: Best-Practice-Empfehlungen für Hochschulen. Zeitschrift für Hochschulentwicklung, 7(3), 1–14. https://doi.org/10.3217/zfhe-7-03/01
Items of the Scale
Antwortformat / Response Scale:
7-stufiges Antwortformat mit den Optionen 1 = “stimme gar nicht zu”, 2 = “stimme nicht zu”, 3 = “stimme eher nicht zu”, 4 = “neutral”, 5 = “stimme eher zu”, 6 = “stimme zu” und 7 = “stimme vollkommen zu”. Zusätzlich steht die Antwortoption “nicht sinnvoll beantwortbar” zur Verfügung.
- Die zu bearbeitenden Basistexte waren verständlich.
- Die Basistexte hatten einen klaren Bezug zu den Zielen der Veranstaltung.
- Die Bearbeitung der Basistexte hat mir beim Verständnis der Lehrinhalte geholfen.
- Der Zeitaufwand für die Bearbeitung der Basistexte war unangemessen hoch.
- Insgesamt bewerte ich die Basistexte als gut.
Hinweis zur Auswertung: Item 4 ist negativ formuliert und muss vor der Skalenwertberechnung umkodiert werden (neuer Wert = 8 – Rohwert). Anschließend können die Items aufsummiert oder gemittelt werden.