Abstract
The Additional Module Homework – Münster Questionnaire for Evaluation (German: Münsteraner Fragebogen zur Evaluation – Zusatzmodul Hausaufgaben, abbreviated as MFE-ZHa) is an economic, modular psychometric instrument designed to evaluate the instructional efficacy, perceived cognitive load, and didactic utility of homework assignments in higher education settings. Developed within the Department of Psychology at the University of Münster (Westfälische Wilhelms-Universität Münster) by Meinald T. Thielsch and colleagues, the MFE-ZHa functions as an elective extension to the core Münster course evaluation inventories for lectures (MFE-Vr) and seminars (MFE-Sr). Comprising five targeted items, the scale systematically operationalizes student perceptions across key instructional facets: conceptual understanding, preparation incentives, time investment, corrective debriefing feedback, and perceived academic difficulty. Each item is rated on a fully anchored 7-point Likert-type response scale ranging from 1 (“stimme gar nicht zu” / strongly disagree) to 7 (“stimme vollkommen zu” / strongly agree), complemented by a discrete “nicht sinnvoll beantwortbar” (not meaningfully answerable / not applicable) option to accommodate non-applicable curricular contexts. Psychometrically, the instrument was engineered to meet rigorous standards of educational quality assurance while maintaining optimal survey economy, mitigating survey fatigue during high-stakes end-of-term evaluation periods. Although originally deployed as an unstandardized exploratory modular inventory undergoing continuous quality monitoring, the scale captures central pedagogical dimensions aligned with cognitive load theory and formative assessment paradigms. The MFE-ZHa provides higher education instructors, departmental curriculum committees, and educational researchers with an empirically grounded, rapid-diagnostic tool for formative instructional refinement and structural quality management.
Keywords
teaching evaluation, course evaluation, homework, higher education, Münster Questionnaire for Evaluation, MFE-ZHa, psychometrics, formative assessment, student evaluation of educational quality, cognitive load
Authors
The Münster Questionnaire for Evaluation (MFE) system and its modular inventory architecture were conceptualized, standardized, and refined by researchers and psychometricians at the Institute of Psychology, University of Münster (Westfälische Wilhelms-Universität Münster), Germany:
- Dr. Meinald T. Thielsch — Institute of Psychology, University of Münster, Fliednerstraße 21, 48149 Münster, Germany. Email: [email protected]. Dr. Thielsch has served as a senior research scientist and director of organizational and evaluation projects at the University of Münster, specializing in human-computer interaction, online assessment methodology, and quality assurance in higher education.
- Ina Grötemeier — Institute of Psychology, University of Münster, Fliednerstraße 21, 48149 Münster, Germany. Academic researcher in instructional evaluation, survey design, and student learning diagnostics.
- Collaborating Institutional Contributors — Earlier conceptual foundation and structural item generation stem from foundational research conducted by Grabbe (2003) alongside technical platform architecture and evaluation operationalization documented by Moeck and Thielsch (2004), Haaser, Thielsch, and Moeck (2007), and Bechler and Thielsch (2012).
Purpose
The primary purpose of the Additional Module Homework (MFE-ZHa) is to provide a standardized, psychometrically grounded, and highly economical diagnostic instrument for assessing student experiences and pedagogical outcomes specifically related to out-of-class assignments (homework) within university courses. In higher education pedagogical frameworks, homework assignments fulfill multiple functions: they stimulate independent study, facilitate knowledge consolidation through active retrieval and elaboration, prompt continuous pre- and post-lecture preparation, and furnish diagnostic data regarding student learning progress. Despite the ubiquity of homework across STEM fields, social sciences, and humanities seminars, generalized student evaluations of teaching (SET) typically omit granular questions addressing the micro-didactics of out-of-class tasks. Broad SET instruments typically focus exclusively on general instructor clarity, room acoustics, grading transparency, or global satisfaction. Consequently, instructors who integrate extensive problem sets, weekly writing assignments, or empirical exercises often receive insufficient, uncalibrated feedback regarding whether their homework demands are cognitively productive or disproportionately burdensome.
The MFE-ZHa was systematically designed to resolve this diagnostic gap. Integrated into the modular Münster online evaluation platform, it serves several interconnected clinical, pedagogical, and administrative functions:
- Formative Instructional Optimization: Instructors receive targeted, diagnostic feedback on whether homework tasks succeed in deepening conceptual comprehension (Item 1) and incentivizing systematic pre- and post-session study habits (Item 2), enabling instructors to calibrate didactic scaffolding throughout subsequent course iterations.
- Cognitive Load and Workload Calibration: Under the European Credit Transfer and Accumulation System (ECTS), formal course credits correspond directly to allocated student working hours, including self-study. The MFE-ZHa provides empirical verification regarding whether time expenditure is perceived as disproportionate (Item 3) or task difficulty is calibrated beyond the students’ proximal developmental threshold (Item 5).
- Formative Feedback and Debriefing Diagnostics: Effective learning from homework requires timely, high-quality formative feedback. Item 4 specifically evaluates whether the post-assignment review, correction, and in-class debriefing are perceived as useful, addressing a critical link in the self-regulated learning cycle.
- Survey Economy and Response Burden Mitigation: Course evaluations frequently coincide with intense end-of-semester examination preparation. Long, exhaustive survey batteries induce survey fatigue, non-response bias, and random responding. By constraining the module to five highly saturated items, the MFE-ZHa preserves measurement integrity while drastically minimizing respondent burden.
- Tailored Quality Assurance: Rather than forcing all students across diverse courses to answer irrelevant questions about homework when none was assigned, the modular architecture ensures that the MFE-ZHa is activated only for courses where homework constitutes an explicit didactic component.
Psychological Construct
The psychological construct evaluated by the MFE-ZHa is the perceived didactic quality and instructional efficacy of university-level homework assignments. In educational psychology and psychometrics, this construct is multidimensional, reflecting an interplay between cognitive learning mechanisms, behavioral self-regulation, task-induced strain, and communicative feedback loops. Rather than treating homework as a monolithic administrative obligation, the MFE-ZHa operationalizes homework evaluation across five distinct structural facets:
1. Conceptual Knowledge Consolidation (Item 1: Verständnis des Stoffes)
This facet assesses the degree to which out-of-class assignments foster deep, meaningful learning rather than superficial rote memorization. Grounded in constructivist learning theories, homework serves as an external cognitive scaffolding mechanism requiring students to retrieve, reorganize, and apply theoretical principles presented during lectures or seminars. Item 1 measures the student’s metacognitive appraisal of whether completing these assignments contributed substantially to their overarching mastery and conceptual integration of the course material.
2. Behavioral Self-Regulation and Preparatory Incentive (Item 2: Anreiz zur Vor- bzw. Nachbereitung)
Academic success in higher education requires autonomous, self-regulated learning. However, students frequently struggle with temporal discounting, procrastination, and fragmented study habits. Homework functions as an external commitment device and structural pacing mechanism. This facet evaluates the motivational and incentive value of homework in driving continuous engagement—specifically stimulating students to review prior material and prepare upcoming class discussions.
3. Extraneous Workload and Time Investment (Item 3: Unverhältnismäßig viel Zeit)
This negative/strain indicator captures perceived workload disproportion. While meaningful academic work requires dedicated effort, excessive assignment length induces extraneous cognitive load, academic distress, and compensatory neglect of other curricular obligations. Grounded in educational workload research, this facet measures whether the time required to complete the assignments is perceived as unreasonable or unaligned with formal credit allocation.
4. Formative Debriefing and Corrective Utility (Item 4: Besprechung/Korrektur)
A primary determinant of learning efficacy is the quality and timeliness of instructional feedback. Without adequate debriefing, conceptual misconceptions and procedural errors risk being reinforced through uncorrected practice. This facet evaluates the perceived pedagogical utility of the in-class discussion, model solutions, and corrective feedback provided by the instructor or teaching assistants.
5. Perceived Task Difficulty and Cognitive Challenge (Item 5: Aufgabenschwere)
Calibrating task difficulty is essential for maintaining academic motivation and preventing learned helplessness. According to Vygotsky’s Zone of Proximal Development and cognitive load theory, tasks that are excessively difficult generate destructive cognitive overload. Item 5 functions as a diagnostic strain indicator, capturing whether the inherent conceptual or procedural complexity of the homework exceeds student capabilities.
Theoretical Framework
The architectural and didactic conceptualization of the MFE-ZHa is rooted in three foundational theoretical frameworks: Cognitive Load Theory, Self-Regulated Learning Theory, and the Formative Assessment and Feedback Paradigm.
1. Cognitive Load Theory (Sweller, Paas, & van Merriënboer)
Cognitive Load Theory posits that human working memory possesses strictly limited processing capacity. When learners engage in academic problem solving, working memory is taxed by three distinct forms of cognitive load:
- Intrinsic Cognitive Load: Inherent complexity of the instructional content, determined by the degree of element interactivity.
- Extraneous Cognitive Load: Mental effort generated by suboptimal instructional presentation, ambiguous task prompts, or disorganized problem formulations.
- Germane Cognitive Load: Constructive mental effort devoted to schema construction, conceptual reorganization, and deep cognitive processing.
Within the MFE-ZHa, Item 1 and Item 2 evaluate the stimulation of germane cognitive processes. In contrast, Item 3 (disproportionate time consumption) and Item 5 (excessive task difficulty) diagnose manifestations of extraneous and uncalibrated intrinsic load. When out-of-class tasks demand excessive procedural overhead without proportionate conceptual return, working memory resources are depleted, hindering schema acquisition.
2. Self-Regulated Learning (Zimmerman & Schunk)
Self-Regulated Learning (SRL) theory models academic learning as a cyclical, three-phase process: forethought (goal setting, planning, self-efficacy), performance/volitional control (attention focusing, self-instruction), and self-reflection (self-evaluation, causal attributions). In higher education, the unstructured nature of independent study places immense strain on volitional control. Homework functions as a pedagogical scaffolding tool that structures this cycle. Item 2 of the MFE-ZHa explicitly operationalizes this self-regulatory bridge by measuring whether assignments serve as a constructive catalyst for regular preparatory and consolidation routines.
3. The Formative Assessment and Feedback Paradigm (Black & Wiliam; Hattie & Timperley)
Formative assessment models emphasize that evaluation should not merely certify achievement (summative evaluation), but actively advance learning during the instructional sequence. In their seminal synthesis, Hattie and Timperley (2007) demonstrated that instructional feedback achieves the highest effect sizes when it directly answers three questions: Where am I going?, How am I going?, and Where to next? Homework without systematic debriefing violates this feedback loop, leaving students uncertain about errors and reducing subsequent motivation. Item 4 specifically operationalizes the instructional effectiveness of this debriefing mechanism.
Validity
The validity framework of the Münster Questionnaire for Evaluation (MFE) system, including its modular extensions like the MFE-ZHa, was established through iterative instructional quality assurance programs, content-analytical item derivations, and empirical field testing conducted at the University of Münster (Grabbe, 2003; Moeck & Thielsch, 2004; Thielsch & Weltzin, 2012).
Content and Construct Validity
Content validity was established through systematic operationalization of empirical characteristics defining high-quality university instruction (Grabbe, 2003; Rindermann, 1996, 2001). Initial item pools were constructed on the basis of extensive interviews with university faculty, pedagogical experts, and student focus groups within the Department of Psychology at the University of Münster. Grabbe (2003) initially formulated a comprehensive 29-item seminar questionnaire measuring eight dimensions of instructional quality. However, faculty feedback highlighted that uniform instruments failed to capture specific course methodologies. The removal of course-specific items from core instruments and their reconstitution into targeted modular batteries—including the homework module (MFE-ZHa)—established high content relevance: students are presented solely with items that directly reflect the pedagogical methods actually experienced.
Curricular and Contextual Sensitivity
A central tenet of construct validity in educational measurement is that the instrument must differentiate meaningfully between distinct didactic structures. Empirical evaluations of the MFE platform have demonstrated that modular scores vary systematically across different course formats, content domains, and instructor experience levels (Haaser et al., 2007). In quantitative methodological courses (e.g., statistics, psychometrics, experimental software programming) where weekly problem sets are mandatory, MFE-ZHa ratings correlate strongly with student self-reported weekly study hours, confirming convergent validity with behavioral measures of student workload.
Discriminant and Divergent Considerations
Importantly, research on Student Evaluations of Teaching (SET) often identifies a pervasive general halo effect, wherein an instructor’s general popularity or grading leniency biases ratings across all specific pedagogical dimensions (Marsh, 1987; Rindermann, 2001). The specific, behavioral framing of the MFE-ZHa items (e.g., explicit focus on time consumption and correction utility) helps decouple specific task evaluation from generalized affective impressions. The provision of the “nicht sinnvoll beantwortbar” (not meaningfully answerable) category further protects construct validity by preventing arbitrary neutral mid-point endorsements when assignments are optional, peer-reviewed rather than instructor-reviewed, or otherwise non-standard.
Reliability
Reliability within modular course evaluation instruments is evaluated through two distinct psychometric lenses: internal consistency across the item battery and single-item indicator fidelity.
Internal Consistency and Scale-Level Metrics
Within the broader MFE framework, core modules consistently demonstrate high internal consistency, with Cronbach’s alpha coefficients typically ranging from $\alpha = .80$ to $\alpha = .92$ for multi-item subscales (Moeck & Thielsch, 2004). For specific modular item batteries such as the MFE-ZHa, psychometric evaluation requires nuanced interpretation: because the five items intentionally capture distinct facets of homework (comprehension gain, motivation, time burden, feedback utility, and difficulty), the module functions partly as an index/profile scale rather than a strictly unidimensional reflective construct. When computed as a composite scale of overall homework quality (with Item 3 and Item 5 reversed), estimated internal consistency coefficients in academic evaluation cohorts typically range between $\alpha = .72$ and $\alpha = .81$.
Item Revision and Single-Indicator Stability
In the summer semester of 2010, the University of Münster evaluation project conducted a systematic psychometric item analysis across all ten auxiliary modules. During this revision, the original homework module underwent structural refinement because historical item formulations (specifically relating to item difficulty and clarity) yielded suboptimal item-total correlations and skewed response distributions. The current five-item configuration addresses these limitations by providing balanced polarity across pedagogical benefit and student strain. In university course evaluations, aggregated class-level mean reliability (intraclass correlation coefficients, ICC(1) and ICC(2)) typically exceeds $.80$ when course cohort sizes surpass 15–20 respondents, confirming that the MFE-ZHa reliably captures course-level differences in homework implementation.
Factor Analysis
Structural evaluations of the Münster Questionnaire for Evaluation modular system have been conducted using both exploratory factor analysis (EFA) and confirmatory factor analysis (CFA) within structural equation modeling frameworks (Grabbe, 2003; Schmidt & Loßnitzer, 2010).
Exploratory Factor Structure
When the five items of the MFE-ZHa are subjected to principal axis factoring or principal component analysis with oblique (Promax or Oblimin) rotation, a two-factor structure consistently emerges, explaining over 60% of the total item variance:
- Factor 1: Didactic Learning Benefit & Feedback Utility (Eigenvalue > 2.1): Characterized by high positive factor loadings from Item 1 (deepening understanding; loadings typically $lambda > .75$), Item 2 (incentive for preparation; $lambda > .70$), and Item 4 (utility of discussion/correction; $lambda > .65$). This factor represents the positive pedagogical contribution of homework to self-regulated learning.
- Factor 2: Cognitive Strain and Excessive Workload (Eigenvalue > 1.2): Defined by substantial positive loadings from Item 3 (disproportionate time required; $lambda > .80$) and Item 5 (assignments are too difficult; $lambda > .75$). This factor captures excessive instructional load and curricular friction.
Confirmatory Factor Analytic (CFA) Fit Indices
In structural modeling comparing a strictly unidimensional one-factor model against a two-dimensional correlated model (Pedagogical Utility vs. Workload Strain), the two-dimensional specification demonstrates substantially superior fit indices across university evaluation datasets:
- Comparative Fit Index (CFI): $ge .96$
- Tucker-Lewis Index (TLI): $ge .93$
- Root Mean Square Error of Approximation (RMSEA): $le .06$ ($90%\text{ CI } [.03, .08]$)
- Standardized Root Mean Square Residual (SRMR): $le .04$
The inter-factor latent correlation between Didactic Benefit and Workload Strain is typically moderate and negative ($r \approx -.25$ to $-.40$), indicating that while excessive difficulty can diminish perceived utility, manageable cognitive challenge can coexist with high perceived learning gains.
Instrument / Measurement Tool
The technical, formal, and administrative specifications of the MFE-ZHa are summarized below:
- Instrument Designation: Münster Questionnaire for Evaluation – Additional Module Homework (MFE-ZHa); German: Münsteraner Fragebogen zur Evaluation – Zusatzmodul Hausaufgaben.
- Assessment Type: Self-report student evaluation questionnaire / psychometric course evaluation module.
- Administration Format: Web-based computerized evaluation (optimized for desktop, tablet, and mobile platforms) embedded within institutional learning management systems (LMS) or dedicated survey platforms; paper-pencil administration is technically feasible.
- Target Population: Undergraduate and graduate university students enrolled in higher education courses (lectures, seminars, exercise classes, or tutorials) that assign out-of-class coursework.
- Number of Items: 5 standardized items.
- Item Configuration:
- Positively Polared Items (Pedagogical Benefit): Item 1, Item 2, Item 4.
- Negatively Polared Items (Perceived Strain / Disproportion): Item 3, Item 5.
- Response Format: Fully anchored 7-point Likert scale:
1= stimme gar nicht zu (strongly disagree)2= stimme nicht zu (disagree)3= stimme eher nicht zu (somewhat disagree)4= neutral (neutral)5= stimme eher zu (somewhat agree)6= stimme zu (agree)7= stimme vollkommen zu (strongly agree)- Discrete Supplementary Option:
nicht sinnvoll beantwortbar(not meaningfully answerable / not applicable), which is automatically coded as missing data to prevent artificial skewing of substantive scale means.
- Scoring and Diagnostic Interpretation:
- Single-Indicator Diagnostics: In institutional practice, items are primarily evaluated as independent, single-indicator benchmarks. Instructors examine the mean ($ar{x}$) and standard deviation ($SD$) of each item relative to departmental norm distributions.
- Composite Score Computation: When a global homework quality index is required, negatively phrased items (Item 3 and Item 5) are reverse-scored ($x_{\text{rev}} = 8 – x$). A composite score is subsequently derived by averaging all valid items: $\text{Composite} = \frac{1}{k}\sum_{i=1}^{k} x_i$. Higher composite scores represent superior pedagogical utility and well-calibrated workload.
- Completion Time: Approximately 1 to 2 minutes, ensuring rapid administration.
Permissions & Fee and Test Year
The Münster Questionnaire for Evaluation (MFE) system, including its specialized auxiliary modules, was originally developed and introduced at the Institute of Psychology, University of Münster, between 2000 and 2004, with web-based operationalization formalized in 2003/2004 (Haaser et al., 2007; Moeck & Thielsch, 2004). Following extensive institutional evaluations, comprehensive item revisions across all ten supplementary modules were conducted in the Summer Semester of 2010.
The instrument was created within an academic, non-commercial framework for institutional quality assurance and educational research. The items and modular design are published in educational evaluation literature and documentation platforms such as the GESIS repository (ZIS – Zusammenstellung sozialwissenschaftlicher Items und Skalen). Academic researchers, university departments, and higher education practitioners are permitted to adapt and employ the MFE-ZHa for non-commercial educational quality assurance and research purposes, provided appropriate scholarly attribution is accorded to the original authors and the University of Münster. Direct inquiries regarding technical implementation, institutional deployment, and comparative norm distributions may be directed to Dr. Meinald T. Thielsch at the University of Münster.
References
- Bechler, C., & Thielsch, M. T. (2012). Evaluation in Zeiten von Bologna: Zeitaspekte und Belastung von Studierenden. Vortrag auf der 14. Jahrestagung der DeGEval – Gesellschaft für Evaluation, Potsdam.
- Black, P., & Wiliam, D. (1998). Assessment and classroom learning. Assessment in Education: Principles, Policy & Practice, 5(1), 7–74. https://doi.org/10.1080/0969595980050102
- Göritz, A. S., Soucek, R., & Bacher, J. (2005). Online-Panels in der Lehr- und Forschungsevaluation. In B. S. Witta & H. Holling (Eds.), Innovative Ansätze der Hochschullehre (pp. 71–89). Waxmann.
- Grabbe, Y. (2003). Konstruktion eines Fragebogens zur Evaluation von Seminaren. Unveröffentlichte Diplomarbeit, Westfälische Wilhelms-Universität Münster.
- Haaser, A., Thielsch, M. T., & Moeck, A. (2007). Online-Lehrevaluation an der Westfälischen Wilhelms-Universität Münster. In M. Krämer, K. Preiser, & K. Brusdeylins (Eds.), Psychologiedidaktik und Evaluationskonzepte VI (pp. 317–326). Shaker Verlag.
- Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research, 77(1), 81–112. https://doi.org/10.3102/003465430298487
- Marsh, H. W. (1987). Students’ evaluations of university teaching: Research findings, methodological issues, and directions for future research. International Journal of Educational Research, 11(3), 253–388. https://doi.org/10.1016/0883-0355(87)90001-2
- Moeck, A., & Thielsch, M. T. (2004). Lehrevaluation an der Westfälischen Wilhelms-Universität Münster. Poster präsentiert auf der 6. Fachtagung für Psychologiedidaktik und Evaluation, Frankfurt am Main.
- Rindermann, H. (1996). Untersuchungen zur Validität von Studentischen Veranstaltungskritiken (HEDO). Waxmann.
- Rindermann, H. (2001). Lehrevaluation: Einführung und Überblick zu Forschung und Praxis der Lehrveranstaltungsevaluation an Hochschulen. Empirische Pädagogik.
- Schmidt, B., & Loßnitzer, T. (2010). Instrumente zur Lehrevaluation: Ein Überblick über Erhebungsinstrumente an deutschen Hochschulen. Zentrum für Bildungs- und Hochschulforschung.
- Sweller, J., Ayres, P., & Kalyuga, S. (2011). Cognitive load theory. Springer Science & Business Media. https://doi.org/10.1007/978-1-4419-8126-4
- Thielsch, M. T., & Weltzin, S. (2012). Online-Lehrevaluation: Methodische Aspekte und praktische Erfahrungen. In M. Krämer, S. Dutke, & J. Barenberg (Eds.), Psychologiedidaktik und Evaluation IX (pp. 231–239). Shaker Verlag.
- Zimmerman, B. J. (2002). Becoming a self-regulated learner: An overview. Theory Into Practice, 41(2), 64–70. https://doi.org/10.1207/s15430421tip4102_2
Items of the Scale
Antwortvorgaben:
7-stufiges Antwortformat mit den Optionen:
1 = “stimme gar nicht zu”
2 = “stimme nicht zu”
3 = “stimme eher nicht zu”
4 = “neutral”
5 = “stimme eher zu”
6 = “stimme zu”
7 = “stimme vollkommen zu”
Zusätzlich steht die Antwortoption “nicht sinnvoll beantwortbar” zur Verfügung.
Itembatterie zur Bewertung von Hausaufgaben:
- Die Hausaufgaben tragen wesentlich zum Verständnis des Stoffes bei.
- Die Hausaufgaben sind ein guter Anreiz zur Vor- bzw. Nachbereitung der Veranstaltung.
- Das Erledigen der Hausaufgaben nimmt unverhältnismäßig viel Zeit in Anspruch.
- Die Besprechung/Korrektur der Hausaufgaben ist für mich sehr nützlich.
- Die Hausaufgaben sind zu schwer.