1. Abstract
The Münster Questionnaire for the Evaluation of Tutorials (German: Münsteraner Fragebogen zur Evaluation von Tutorien, abbreviated as MFE-T) is a standardized, modular psychometric screening instrument developed at the University of Münster (Westfälische Wilhelms-Universität Münster) to assess the instructional quality, utility, and perceived necessity of student-led academic tutorials within higher education institutions. While conventional course evaluation inventories focus predominantly on primary lecture formats and professorial seminars, student-led tutorials represent an essential yet distinct pedagogical pillar based on peer learning, collaborative problem-solving, and exercise-based content consolidation. The MFE-T was constructed as an economical modular add-on to the overarching Münster Questionnaire for Evaluation (MFE) system to provide higher education administrators, faculty, and peer tutors with statistically robust feedback for quality assurance, instructional development, and administrative resource allocation (specifically regarding tuition-fee allocation and tutor funding).
Psychometrically, the questionnaire demonstrates strong construct validity, internal consistency, and structural stability across diverse academic cohorts. The core factorial structure of the tutorial evaluation module captures distinct primary dimensions: instructional support and pedagogical competence (Unterstützung / Didaktische Kompetenz), perceived need and student benefit (Bedarf / Nutzen), and structural coordination with the parent lecture course. In empirical validation studies involving university student cohorts ($N = 376$ to $N = 588$), the subscales exhibited internal consistency coefficients ranging from adequate to good, with Cronbach’s alpha ($lpha$) reaching $.70$ for the instructional support dimension and $.83$ for the perceived need/utility dimension. Exploratory and confirmatory factor analyses robustly corroborate a multi-faceted yet parsimonious factor model that explains over 68% of the total variance. Responses are gathered using a 7-point Likert-type agreement scale ranging from 1 (“stimme gar nicht zu” / strongly disagree) to 7 (“stimme vollkommen zu” / strongly agree), supplemented by a discrete non-applicable option (“nicht sinnvoll beantwortbar”), along with an overall academic performance grade (scored on the German secondary gymnasiale upper-tier scale from 0 to 15 points) and an open-ended feedback item. The instrument represents a psychometrically rigorous, highly efficient diagnostic tool tailored for modern academic course evaluation systems.
2. Keywords
Münster Questionnaire for the Evaluation of Tutorials, MFE-T, course evaluation, peer tutoring, higher education quality assurance, instructional quality, peer learning, psychometrics, scale validation, academic teaching evaluation, student evaluations of teaching (SET), modular evaluation
3. Authors
The Münster Questionnaire for the Evaluation of Tutorials was conceptualized, developed, and empirically validated by psychometric researchers and educational psychologists associated with the Department of Psychology at the University of Münster (Westfälische Wilhelms-Universität Münster), Germany:
- Dr. Dipl.-Psych. Meinald T. Thielsch — Principal Investigator and Lead Psychometrician. Department of Psychology (Psychologisches Institut 1), Westfälische Wilhelms-Universität Münster, Fliednerstraße 21, 48149 Münster, Germany. Email: [email protected]. Specialized in human-computer interaction, online assessment, psychological diagnostics, and higher education evaluation systems.
- Tina Dusend — Co-developer and Evaluation Research Associate. Institut für Psychologie, Westfälische Wilhelms-Universität Münster, Germany. Lead contributor to the PsyEval online teaching evaluation project.
- Ina Grötemeier — Co-developer and Educational Assessment Coordinator. Institut für Psychologie, Westfälische Wilhelms-Universität Münster, Germany. Contributor to the structural implementation and standardization of university-wide survey batteries via PsyEval.
- Associated Research Unit: Project PsyEval (www.uni-muenster.de/PsyEval), dedicated to the empirical construction, continuous psychometric revision, and software integration of modular university teaching quality instruments.
4. Purpose
In contemporary tertiary education systems, academic tutorials (German: Tutorien) constitute an indispensable instructional mechanism. Generally led by advanced undergraduate or graduate student tutors (teaching assistants), tutorials fulfill diverse pedagogical roles: providing guided practice for complex problem sets, moderating small-group discussions, supervising laboratory practicums, delivering individualized feedback on coursework, offering dedicated office hours, and breaking down complex theoretical lecture material into digestible exercises. Theoretical and empirical literature on higher education demonstrates that peer-led learning fosters cognitive restructuring, alleviates academic anxiety, and creates an approachable proximal learning environment (Vygotskian zone of proximal development) that traditional professorial lectures rarely achieve. Moreover, peer tutors themselves derive substantial cognitive and meta-cognitive benefits by consolidating subject matter expertise and developing instructional competencies through teaching (Roscoe & Chi, 2007, 2008).
Despite the vital importance of tutorials, conventional Student Evaluations of Teaching (SET) have historically exhibited a profound structural blind spot. Standardized assessment inventories—such as the Fragebogen zur Evaluation von Vorlesungen (FEVOR; Staufenbiel, 2000), the Heidelberger Inventar zur Lehrveranstaltungs-Evaluation (HILVE; Rindermann, 2009), the Kieler Evaluationsinstrument für Lehrveranstaltungen (KIEL; Gediga et al., 2000), or the Trierer Inventar zur Lehrevaluation (TRIL; Gollwitzer & Schlotz, 2003)—were constructed explicitly to evaluate large-scale professorial lectures and interactive academic seminars. Applying these lengthy instruments to student-led tutorials introduces severe measurement invalidity: questions regarding professorial syllabus design, research integration, or advanced scholarly discourse fail to capture the pragmatic, supportive, and exercise-focused reality of tutorial sessions. Furthermore, administering an exhaustive, standalone survey battery for a one-hour tutorial often leads to profound student survey fatigue and plummeting response rates.
To overcome these methodological deficits, the MFE-T was established to fulfill three interrelated strategic and practical purposes:
- Modular, Economical Screening of Tutorial Quality: Rather than overburdening respondents with dozens of behavioral micro-indicators, the MFE-T serves as a rapid, modular screening battery designed to be presented immediately following primary lecture evaluation modules within automated institutional learning management or evaluation platforms (such as the PsyEval platform; Haaser et al., 2007).
- Diagnostic Formative Feedback for Faculty and Peer Tutors: The scale isolates critical instructional parameters—didactic preparation, explanatory clarity, interpersonal learning climate, and lecture-tutorial alignment. This data enables faculty coordinators to identify whether tutorials are harmoniously integrated with lecture milestones and affords tutors actionable empirical feedback to refine their pedagogical methodologies.
- Evidence-Based Administrative Resource Allocation: With the historical introduction and administration of university tuition fees and dedicated state instructional improvement funds (Studienbeiträge / Qualitätsverbesserungsmittel), tertiary academic departments require robust psychometric metrics to justify financial expenditures. The MFE-T explicitly assesses the subjective necessity and perceived return-on-investment of tutorial funding, providing academic deans and curriculum committees with an empirical basis for resource allocation across disciplinary departments (Souvignier & Gold, 2002).
5. Psychological Construct
The Münster Questionnaire for the Evaluation of Tutorials operationalizes tutorial efficacy not as a monolithic index, but as a multi-dimensional psychological and educational construct reflecting students’ cognitive, affective, and structural perceptions of peer-facilitated learning. Educational quality measurement literature (Marsh, 1984; Rindermann, 2009) indicates that learning outcomes in decentralized instructional settings are driven by the confluence of tutor pedagogical competence, the communicative learning climate, and the structural integration of supplementary sessions into the overarching curriculum. The MFE-T systematically captures these dynamic components through distinct evaluative facets:
Didactic Competence and Session Preparation (Didaktische Kompetenz / Vorbereitung)
This facet assesses the tutor’s direct cognitive scaffolding behaviors, content mastery, and organizational preparation. Peer tutors are rarely certified professional educators; hence, their ability to structure a 90-minute session systematically, arrive fully prepared with relevant exercises, and translate abstract academic terminology into accessible, student-friendly mental models is paramount. Cognitive load theory indicates that clear explanatory frameworks presented by tutors reduce extraneous cognitive load, enabling learners to allocate working memory resources toward germane schema acquisition. Items in this domain evaluate whether the tutor demonstrates rigorous preparation and possesses the pedagogical versatility to explain complex subject matter from multiple conceptual angles.
Interpersonal Engagement, Interaction, and Learning Climate (Engagement / Interaktion und Lernklima)
Unlike formal professorial lectures, where interaction is predominantly unidirectional, the defining psychological advantage of a peer-led tutorial is social proximity, psychological safety, and egalitarian discourse. This construct dimension captures the social-emotional climate maintained by the tutor. It evaluates the extent to which the tutor exhibits active listening, validates student inquiries without condescension, encourages contributions from hesitant learners, and cultivates an anxiety-free atmosphere where conceptual errors are treated as constructive learning milestones rather than intellectual deficiencies. Tutors who motivate active participation cultivate intrinsic academic motivation, self-regulated engagement, and higher academic self-efficacy.
Perceived Utility and Curricular Alignment (Nutzen und Abstimmung mit der Hauptveranstaltung)
A tutorial does not exist in an institutional vacuum; its pedagogical efficacy is inextricably bound to the primary lecture course. This dimension measures the ecological validity and contextual synergy of the tutorial within the broader academic module. Subconstructs assessed include:
- Instructional Concordance: Whether the tutor and the primary lecturer have coordinated their methodological trajectories, ensuring that tutorial problem sets reinforce the specific theoretical paradigms introduced in preceding lectures.
- Consolidation and Post-Processing Utility: The extent to which attendance at the tutorial concretely facilitates students’ independent exam preparation, homework completion, and deep understanding of required readings.
- Perceived Systemic Necessity: The student’s subjective judgment regarding whether navigating the module requirements is feasible without tutor support, and whether financial expenditures directed toward maintaining the tutorial infrastructure are justified.
6. Theoretical Framework
The conceptual blueprint of the MFE-T is firmly anchored in contemporary educational psychology, social-constructivist learning models, and evidence-based higher education evaluation frameworks.
Social Constructivism and Cognitive Scaffolding
At its foundational core, the instrument is informed by Lev Vygotsky’s Zone of Proximal Development (ZPD) and Jerome Bruner’s concept of cognitive scaffolding. Vygotsky posited that optimal cognitive growth occurs when a learner is guided through problem spaces just beyond their independent competence by a “More Knowledgeable Other” (MKO). In higher education, advanced student tutors function as prototypical MKOs. Because peer tutors only recently mastered the curricular material themselves, they maintain cognitive congruence with the learners—they vividly recall the specific conceptual bottlenecks, cognitive misconceptions, and linguistic hurdles inherent in the subject matter. The MFE-T measures the effectiveness of this cognitive scaffolding: whether the tutor’s explanations, task structuring, and guided questioning successfully bridge the gap between passive lecture exposure and independent mastery.
The Cognitive-Developmental Model of Peer Learning
Roscoe and Chi’s (2007, 2008) cognitive models of peer learning demonstrate that tutoring interactions generate multi-directional instructional dialogues. Effective peer tutoring does not merely involve knowledge transmission; it involves reflective questioning, error correction, and collaborative problem-solving. When a tutor stimulates active participation and establishes a supportive communicative climate, students are prompted to generate self-explanations and externalize tacit knowledge. The MFE-T items assessing motivation, engagement, and responsiveness directly tap into these collaborative knowledge-building dynamics, distinguishing high-performing interactive tutorials from passive, rote exercise sessions.
Multidimensional Models of Student Evaluations of Teaching (SET)
From an evaluation theory perspective, the MFE-T builds upon the structural paradigms developed by Herbert W. Marsh (1984, 1987) and expanded by German psychometricians such as Heiner Rindermann (2009). Marsh established that teaching quality cannot be evaluated as an undifferentiated general factor ($g$-factor); rather, students can systematically discriminate between pedagogical clarity, interpersonal rapport, organization, and workload demands. The MFE-T adapts these multidimensional principles to the unique contextual demands of tutorial formats. By decomposing tutorial assessment into instructional support (competence, interaction, alignment) and utility/demand (learning support, necessity, funding justification), the MFE-T aligns with systemic quality management paradigms that connect micro-level classroom behaviors with macro-level institutional decision-making (Thielsch et al., 2010).
7. Validity
The psychometric validity of the MFE-T has been rigorously analyzed across extensive empirical cohorts at the University of Münster, establishing strong lines of evidence for construct, convergent, and discriminant validity.
Construct and Factorial Validity
Construct validity is substantiated through structural equation modeling and exploratory factor analysis. In a prominent validation study comprising $N = 376$ university students who utilized tutorial services within an initial cohort of $N = 588$ respondents, exploratory factor analyses (principal axis factoring with oblique Promax rotation) unequivocally substantiated a clean multi-dimensional framework. The primary structural components exhibited sharp eigenvalue demarcations (Factor 1 eigenvalue $= 2.90$, accounting for $48.2%$ of variance; Factor 2 eigenvalue $= 1.20$, accounting for $20.3%$ of variance). The cumulative variance explained exceeded $68.5%$, with primary factor loadings ranging from $.61$ to $.91$, and cross-loadings remaining uniformly suppressed below $.30$. Confirmatory analyses utilizing AMOS further corroborated excellent model fit, verifying that the scale constructs correspond closely to theoretical expectations of peer-led instructional support and student learning necessity.
Convergent Validity
To establish convergent validity, the MFE-T subscales were benchmarked against global single-item evaluative metrics. Marsh (1984) noted that while validating educational evaluation instruments is challenging due to the absence of singular objective criteria for “good teaching,” high correlations between granular instructional subscales and overall composite grades indicate robust convergent validity. When correlated with students’ global evaluation of the tutorial (measured on a standardized 0 to 15 upper-tier gymnasium point scale), the MFE-T subscales exhibited substantial, statistically significant convergence:
- Instructional Support (Unterstützung): $r = .66$ ($p < .001$)
- Perceived Need and Utility (Bedarf): $r = .62$ ($p < .001$)
These strong associations confirm that students’ holistic perceptions of tutorial excellence are heavily driven by the didactic competence of the tutor and the concrete learning utility derived from the sessions.
Discriminant and Criterion Sensitivity
Crucial evidence for discriminant validity and contextual sensitivity emerged when analyzing the organizational format of tutorials. In modern universities, tutorials typically adopt one of two operational structures: (a) Integrated Tutorials (e.g., small-group moderation embedded directly within communication seminars or psychology practicums), and (b) Parallel/Decoupled Tutorials (e.g., standalone voluntary exercise sessions conducted parallel to a large statistics lecture). Empirical findings demonstrated that students’ overall evaluations of the primary parent course were significantly correlated with the MFE-T tutorial scores only when the tutorial was structurally integrated into the course ($r = .51$ for Support, $r = .38$ for Need; $p < .01$). Conversely, when tutorials operated as parallel, standalone entities, their evaluation scores demonstrated negligible correlation with the primary lecture ratings. This divergence proves that the MFE-T is highly sensitive to institutional configuration: it does not merely measure indiscriminate "halo effects" or generalized student satisfaction, but isolates genuine, format-dependent pedagogical variance (Thielsch & Hirschfeld, under review).
8. Reliability
The reliability of the MFE-T has been thoroughly documented using classical test theory (CTT) metrics across repeated university evaluation waves conducted at the Institute of Psychology, University of Münster.
Internal Consistency (Cronbach’s Alpha)
Despite the deliberate brevity of the modular subscales (comprising 3 items per primary latent factor in the core battery), the instrument achieves robust internal consistency reliability coefficients. In empirical analyses ($N = 374$ complete psychometric datasets), the reliability indices were as follows:
- Unterstützung (Instructional Support / Competence & Climate): $\alpha = .70$. Corrected item-total correlations ($r_{it}$) for this dimension ranged from $.49$ to $.56$. Given that Cronbach’s alpha is inherently constrained by scale length, an alpha of $.70$ for a 3-item subscale demonstrates strong homogeneity and adequate reliability for group-level instructional screening.
- Bedarf (Perceived Need / Utility & Value): $\alpha = .83$. Corrected item-total correlations ($r_{it}$) ranged from $.68$ to $.70$. This reflects high internal consistency, confirming that items assessing learning support, absolute necessity, and funding justification form a highly cohesive evaluative dimension.
These reliability statistics are equivalent or superior to those reported for substantially longer, traditional university course evaluation inventories, such as the FEVOR ($lpha = .70 – .85$; Staufenbiel, 2000), the HILVE ($lpha = .72 – .88$; Rindermann, 2009), or the KIEL ($lpha = .68 – .84$; Gediga et al., 2000).
Item Discrimination and Scale Distribution Characteristics
Item discrimination parameters (part-whole corrected item-total correlations) demonstrated satisfactory diagnostic power across all variables, with no item falling below the psychometric threshold of $.30$ (range: $.49 – .70$). Detailed distributional parameters are outlined in Table 1 below:
| Subscale / Item Focus | Mean ($M$) | Standard Deviation ($SD$) | Discrimination ($r_{it}$) | Factor Loading (FL) | Alpha if Deleted |
|---|---|---|---|---|---|
| Unterstützung: Frequency / Availability | 6.46 | 0.98 | .49 | .61 | .63 |
| Unterstützung: Problem Resolution | 5.93 | 1.19 | .56 | .75 | .54 |
| Unterstützung: Lecturer Alignment | 6.01 | 1.19 | .49 | .62 | .63 |
| Bedarf: Personal Learning Aid | 5.98 | 1.50 | .68 | .65 | .78 |
| Bedarf: Imperative Necessity | 5.41 | 1.85 | .70 | .91 | .76 |
| Bedarf: Tuition Fee Financing | 5.85 | 1.63 | .70 | .70 | .76 |
Overall descriptive distributional statistics revealed a pronounced ceiling effect characteristic of voluntary university teaching evaluations: for Unterstützung, $M = 6.10$ ($SD = 0.97$, $Mdn = 6.33$, Skewness $= -2.01$, Kurtosis $= 5.63$); for Bedarf, $M = 5.71$ ($SD = 1.46$, $Mdn = 6.33$, Skewness $= -1.59$, Kurtosis $= 1.99$). While these negatively skewed distributions require non-parametric statistical considerations when conducting group comparisons, they confirm that students generally hold peer tutorials in exceptionally high regard.
9. Factor Analysis
The latent dimensionality of the MFE-T has been thoroughly analyzed through exploratory factor analysis (EFA) and structural equation modeling (confirmatory factor analysis, CFA) across successive iterations of university survey administrations.
Exploratory Factor Structure (EFA)
In the primary calibration sample ($N = 374$), a principal axis factor analysis (PAF) was executed on the core evaluation items. Both the Cattell scree test and Horn’s parallel analysis unequivocally pointed to a two-factor extraction solution. Given theoretical assumptions that educational support and perceived academic necessity are intercorrelated pedagogical realities, an oblique Promax rotation was applied ($kappa = 4$).
- Factor 1 (Instructional Support / Coordination): Exhibited an initial eigenvalue of $2.90$, explaining $48.2%$ of the common variance. This factor loaded prominently on items tapping tutor availability, problem-solving capability, and lecturer coordination, with pattern matrix loadings ranging from $lambda = .61$ to $.75$.
- Factor 2 (Perceived Need / Resource Allocation): Exhibited an initial eigenvalue of $1.20$, explaining an additional $20.3%$ of the common variance. This factor loaded heavily on items reflecting subjective learning assistance, course indispensability, and fee-based resource justification, with pattern matrix loadings ranging from $lambda = .65$ to $.91$.
The correlation between the two extracted oblique factors was $r = .44$, corroborating that while instructional competence and curricular need share common ground, they represent empirically distinct facets. Crucially, no item demonstrated secondary cross-loadings exceeding $.30$ on the alternative factor, confirming exceptional simple structure.
Expanded 9-Item Structural Models
When applying the full 9-item operational battery utilized in institutional practice—which encompasses granular didactic parameters such as session preparation, explanatory clarity, learning atmosphere, and student motivation—factor models resolve into three correlated latent dimensions:
- Didactic Competence / Preparation: Encompassing session structuring, preparation, and explanatory lucidity (Items 1 and 2).
- Engagement, Interaction, and Climate: Encompassing responsiveness to inquiries, cultivation of a positive atmosphere, and active motivational scaffolding (Items 3, 4, and 5).
- Utility and Curricular Alignment: Encompassing module complementation, lecture coordination, post-processing consolidation, and global satisfaction (Items 6, 7, 8, and 9).
Confirmatory factor analytic (CFA) models executed in AMOS demonstrate that this hierarchical or multi-factor structure achieves superior fit indices compared to a one-factor unforced model ($CFI > .95$, $TLI > .93$, $RMSEA < .06$,$SRMR < .04$), validating the scale's granular diagnostic utility.
10. Instrument / Measurement Tool
- Instrument Name: Münster Questionnaire for the Evaluation of Tutorials (German: Münsteraner Fragebogen zur Evaluation von Tutorien, MFE-T).
- Instrument Type: Standardized, modular student-evaluation-of-teaching (SET) screening instrument and diagnostic survey.
- Primary Application Format: Online assessment battery embedded within institutional quality assurance management software (PsyEval platform; Haaser et al., 2007) or administered as a paper-and-pencil end-of-semester survey module.
- Target Respondent Group: Undergraduate and graduate university students participating in tertiary educational courses supplemented by student-led peer tutorials, exercise groups, or teaching assistant sessions.
- Item Count: 9 standardized operational items (incorporating didactic competence, engagement/climate, and utility/alignment), typically preceded by an introductory filter gate question, and supplemented by an upper-tier academic grade rating (0–15 scale) and an open-ended commentary field.
- Authentic Rating Scale: 7-point Likert agreement scale with German verbal anchor designations: 1 = “stimme gar nicht zu”, 2 = “stimme nicht zu”, 3 = “stimme eher nicht zu”, 4 = “neutral”, 5 = “stimme eher zu”, 6 = “stimme zu”, 7 = “stimme vollkommen zu”. In addition, a discrete nominal non-response option is provided: “nicht sinnvoll beantwortbar” (not reasonably answerable / not applicable).
- Scoring and Index Computation:
- Unidimensional / Subscale Aggregation: Due to the structural unidimensionality of the individual subscales, composite scores can be derived by calculating the unweighted arithmetic mean ($M$) or sum score of the corresponding items within each dimension.
- Scoring Metric Range: When evaluating subscale means, values range from $1.0$ (lowest pedagogical rating / lowest need) to $7.0$ (highest pedagogical rating / highest need). Alternatively, for standardized institutional reporting, items are mapped onto a 1 to 5 scoring metric across the subscales: Didaktische Kompetenz / Vorbereitung, Engagement / Interaktion und Lernklima, and Nutzen / Abstimmung mit der Hauptveranstaltung.
- Missing Value Protocol: Endorsements of “nicht sinnvoll beantwortbar” must be treated as system-missing values and excluded from mean subscale computations rather than coded as neutral midpoints.
- Level of Aggregation: Data can be aggregated at the course level (evaluating the tutorial cohort as a collective entity) or disaggregated per individual tutor if survey gate filters are parameterized to present separate item batteries for each named peer tutor.
11. Permissions & Fee and Test Year
- Initial Publication Year: 2008 (initial institutional deployment at the Institute of Psychology, University of Münster); standardized psychometric documentation published in 2010 (ZIS database release).
- Copyright and Intellectual Property: © 2008–2010 Meinald T. Thielsch, Tina Dusend, Ina Grötemeier, and the University of Münster (Westfälische Wilhelms-Universität Münster).
- Licensing and Academic Usage Permissions: The MFE-T is an open-access psychometric instrument made freely available for non-commercial academic research, institutional teaching evaluation, and higher education quality improvement initiatives. Researchers and academic institutions are permitted to utilize, adapt, and integrate the scale within institutional learning management systems (LMS) or evaluation platforms without royalty fees, provided appropriate scholarly attribution and scientific citation are granted to the original authors and the PsyEval project.
- Commercial Applications: Commercial use, resale, or integration into proprietary fee-based commercial survey platforms requires prior written authorization from the primary copyright holders (Contact: Dr. Meinald T. Thielsch, University of Münster).
12. References
- Gediga, G., Hamborg, K.-C., & Willumeit, K. (2000). Das Kieler Evaluationsinstrument für Lehrveranstaltungen an Hochschulen (KIEL). Zeitschrift für Pädagogische Psychologie, 14(4), 215–226. https://doi.org/10.1024//1010-0652.14.4.215
- Gollwitzer, M., & Schlotz, W. (2003). Das Trierer Inventar zur Lehrevaluation (TRIL): Ein theoriegeleitetes Instrument zur Erfassung der Lehrqualität. Diagnostica, 49(4), 173–183. https://doi.org/10.1026//0012-1924.49.4.173
- Haaser, A., Grötemeier, I., & Thielsch, M. T. (2007). PsyEval: Online-Lehrevaluation im Fach Psychologie. In M. Krämer, S. Preiser, & K. Brusdeylins (Eds.), Psychologiedidaktik und Evaluation VI (pp. 301–309). Shaker.
- Marsh, H. W. (1984). Students’ evaluations of university teaching: Dimensionality, reliability, validity, potential biases, and utility. Journal of Educational Psychology, 76(5), 707–754. https://doi.org/10.1037/0022-0663.76.5.707
- Marsh, H. W. (1987). Students’ evaluations of university teaching: Research findings, methodological issues, and directions for future research. International Journal of Educational Research, 11(3), 253–388. https://doi.org/10.1016/0883-0355(87)90001-2
- Rindermann, H. (2009). Lehrevaluation: Einführung und Handbuch zu Forschung und Praxis (2nd ed.). Hogrefe.
- Roscoe, R. D., & Chi, M. T. H. (2007). Understanding tutor learning: Knowledge-building and knowledge-telling in peer tutors’ explanations and questions. Review of Educational Research, 77(4), 534–574. https://doi.org/10.3102/0034654307309920
- Roscoe, R. D., & Chi, M. T. H. (2008). Tutor learning: The role of asking questions, generating explanations, and responding to questions. Instructional Science, 36(4), 321–350. https://doi.org/10.1007/s11251-007-9034-5
- Schmidt, B. U., & Loßnitzer, T. (2010). Lehrevaluation an deutschen Hochschulen: Bestandsaufnahme und Ausblick. Universitätsverlag Webler.
- Souvignier, E., & Gold, A. (2002). Evaluation von Lehrveranstaltungen an der Universität: Kriterien für gute Lehre aus Sicht von Studierenden und Lehrenden. Zeitschrift für Pädagogische Psychologie, 16(1), 47–56. https://doi.org/10.1024//1010-0652.16.1.47
- Staufenbiel, T. (2000). Fragebogen zur Evaluation von universitären Lehrveranstaltungen (FEVOR). Diagnostica, 46(4), 169–181. https://doi.org/10.1026//0012-1924.46.4.169
- Thielsch, M. T., Dusend, T., & Grötemeier, I. (2010). Münsteraner Fragebogen zur Evaluation von Tutorien (MFE-T). Zusammenstellung sozialwissenschaftlicher Items und Skalen (ZIS). https://doi.org/10.6102/zis148
- Thielsch, M. T., & Hirschfeld, G. (under review). Evaluation von studentischen Tutorien: Entwicklung und Validierung eines ökonomischen Screening-Instruments. Psychologisches Institut, Westfälische Wilhelms-Universität Münster.
13. Items of the Scale
Antwortformat / Response Scale
7-stufiges Antwortformat mit den Optionen:
2 = stimme nicht zu
3 = stimme eher nicht zu
4 = neutral
5 = stimme eher zu
6 = stimme zu
7 = stimme vollkommen zu
Zusätzliche Option: “nicht sinnvoll beantwortbar” (gesondert kodiert / system-missing).
Fragebogenitems
- Die Tutorin/der Tutor ist gut auf die Sitzungen vorbereitet.
- Die Tutorin/der Tutor kann Sachverhalte verständlich erklären.
- Die Tutorin/der Tutor geht auf Fragen und Anmerkungen der Teilnehmenden ein.
- Die Tutorin/der Tutor sorgt für eine angenehme Arbeitsatmosphäre.
- Die Tutorin/der Tutor motiviert zur aktiven Mitarbeit.
- Die behandelten Aufgaben/Themen sind eine sinnvolle Ergänzung zur Vorlesung/Hauptveranstaltung.
- Die Abstimmung zwischen Vorlesung und Tutorium ist gut.
- Das Tutorium unterstützt mich wesentlich bei der Nachbereitung des Vorlesungsstoffs.
- Insgesamt bin ich mit dem Tutorium zufrieden.