Abstract
The Feedback Tool for Responder Development – Senior Staff (FIRE-CU) (German: Feedback zur Rettungskräfteentwicklung – Führungsstab) is a standardized, psychometrically validated multidimensional evaluation instrument designed specifically for training programs directed at rescue forces and fire service command staff members (Führungsstab; “CU” designates “Command Unit”). Originating from a collaborative research initiative between the Department of Organizational and Business Psychology at the University of Münster and the State Fire Service Institute of North Rhine-Westphalia (Institut der Feuerwehr Nordrhein-Westfalen; IdF NRW), the FIRE-CU represents an adaptation and structural expansion of the foundational FIRE questionnaire (Schulte & Thielsch, 2019). The instrument systematically bridges educational process quality and immediate training outcomes across high-stakes emergency management domains.
The core instrument consists of 22 items distributed across six psychometric subscales, conceptually stratified into two higher-order evaluation tiers: the Learning Process Level (comprising Instructor Behavior [Dozentenverhalten], Requirement Level [Anforderungsniveau], Structure [Struktur], and Group Dynamics [Gruppe]) and the Learning Outcome Level (comprising Competence Acquisition [Kompetenzerwerb] and Training Transfer [Transfer]). Four supplemental, non-scale items capture overarching global assessments (Globalurteil) and qualitative participant remarks (Feedback). The core items are assessed using an unforced seven-point Likert scale (1 = stimme gar nicht zu / strongly disagree to 7 = stimme vollkommen zu / strongly agree) supplemented with an explicit non-substantive filter category (nicht sinnvoll beantwortbar / not meaningfully answerable). Structural equation modeling and confirmatory factor analysis (CFA) validate the multi-tier factor structure, confirming acceptable to excellent model fit (Process level: CFI = .97, TLI = .96, RMSEA = .05, SRMR = .04; Outcome level: CFI = .97, TLI = .95, RMSEA = .06, SRMR = .04). Reliability analyses demonstrate strong internal consistency across dimensions, with Cronbach’s alpha spanning .75 to .90 and McDonald’s omega ranging from .76 to .91. The instrument provides emergency service training academies, crisis management organizations, and military institutions with an evidence-based diagnostic protocol for curriculum development, quality assurance, and instructional benchmarking.
Keywords
FIRE-CU, emergency responder training, command staff evaluation, crisis management, High Reliability Organizations (HRO), incident command systems, training evaluation, confirmatory factor analysis, Shared Mental Models, Kirkpatrick model, fire service leadership, psychometrics
Authors
The FIRE-CU was conceptualized, developed, and validated through an institutional research partnership between academic organizational psychologists and operational command staff training directors:
- Meinald T. Thielsch, apl. Prof. Dr. rer. nat. — Department of Psychology, Organizational and Business Psychology (OWMS), University of Münster, Münster, Germany. Academic lead for public safety and emergency service psychometric evaluation instruments.
- Edin Hadzihalilovic, M.Sc. — Department of Psychology, Organizational and Business Psychology (OWMS), University of Münster, Münster, Germany. Specialized in psychometric modeling, human factors in critical incidents, and organizational feedback interventions.
- Cooperation Partner: Institut der Feuerwehr Nordrhein-Westfalen (IdF NRW) — Division of Crisis Management and Research (Abteilung Krisenmanagement und Forschung), Münster, Germany. The largest state fire service training institute in Germany, providing curriculum architecture, expert command panels, and empirical data gathering infrastructure.
Purpose
The operational management of large-scale disasters, major industrial incidents, widespread natural catastrophes, and multi-agency mass casualty events exceeds the coordination capacity of field-level tactical commanders. Within modern incident management frameworks—such as the German Fire Service Regulation 100 (Feuerwehr-Dienstvorschrift 100; FwDV 100) or international equivalents like the Incident Command System (ICS)—strategic decision-making is delegated to an operatively deployed Command Unit or Command Staff (Führungsstab). In Germany, members of these strategic groups are seasoned executive officers who have advanced beyond company-level and battalion-level commanding qualifications (specifically, holding certification as a Verbandsführer, the tactical leader of multiple operational platoons). Because full-scale catastrophic emergencies requiring operational-tactical command staff mobilization occur with low base-rate frequency—historically estimated at once every 25 years per localized jurisdiction, punctuated by severe crises such as pandemic response or regional flooding—staff competence depends almost entirely on high-fidelity simulation and specialized command post exercises (Stabsrahmenübungen).
Despite the critical nature of command staff preparedness, evaluation practices in emergency responder curricula historically relied on unstandardized, intuitive feedback, or general-purpose academic course ratings that fail to capture the socio-cognitive demands of crisis command. The purpose of the FIRE-CU is to provide an empirically grounded, methodologically robust, and actionable evaluation instrument tailored to the structural, instructional, and psychological realities of strategic command staff development. The instrument fulfills three central functions:
- Formative Instructional Diagnostics: By assessing distinct facets of instructional design (e.g., didactic pace, structural lucidity, and instructor feedback), the tool enables trainers and academic leadership to pinpoint pedagogical deficiencies, optimize syllabus pacing, and tailor exercise difficulty to participant capabilities.
- Competence and Transfer Assurance: Unlike superficial reaction sheets, the FIRE-CU explicitly models subjective capability development in high-stakes non-technical skills, including rapid information processing, decision-making under stress, boundary awareness, and real-world transfer readiness.
- Benchmarking Across High-Reliability Contexts: Although developed within the fire service, the scale items were deliberately framed to be non-parochial. This structural generality allows cross-organizational deployment across police command bodies, emergency medical disaster boards, the Federal Agency for Technical Relief (Technisches Hilfswerk; THW), civil protection administrative staffs (Krisenstab / Verwaltungsstab), and military strategic units.
Psychological Construct
The FIRE-CU measures the perceived quality of crisis management training through an integrated, multidimensional architecture. Drawing from training psychology and human factors engineering, the instrument segments educational appraisal into six latent sub-constructs operating across two sequential stages of the training lifecycle: process dynamics and outcome attainment.
1. Learning Process Constructs (Prozessebene)
The process tier captures the environmental, didactic, and interpersonal dynamics that occur during the delivery of instruction. Rather than viewing the learning environment as a monolithic entity, the FIRE-CU isolates four distinct psychological and organizational elements:
- Instructor Behavior (Dozentenverhalten, Items 1–4): Evaluates the pedagogical and interpersonal effectiveness of the faculty. This subscale measures the instructor’s ability to condense complex emergency phenomena into precise operational concepts (Item 1), deliver constructive and actionable behavioral feedback (Item 2), foster active intellectual and tactical contribution (Item 3), and project an authentic commitment to trainee learning success (Item 4). In adult learning within High Reliability Organizations, instructor credibility and motivational support serve as key determinants of psychological safety.
- Requirement Level (Anforderungsniveau, Items 5–7): Measures cognitive load, curricular volume, and instructional tempo. Comprising items assessing overwhelming content breadth (Item 5), excessive pacing (Item 6), and excessive conceptual difficulty (Item 7), this subscale functions as an inverted indicator of curricular appropriateness. In crisis leadership development, an excessive requirement level triggers cognitive overload, whereas an insufficient requirement level fails to elicit the stress inoculation necessary for real-world incidents.
- Structure (Struktur, Items 8–10): Quantifies the conceptual coherence and transparency of the instructional pathway. It addresses structural clarity (Item 8), the longitudinal comprehensibility of the training agenda (Item 9), and the delivery of a comprehensive, integrated domain overview (Item 10). Transparent structure reduces extraneous cognitive load, freeing mental bandwidth for complex tactical problem-solving.
- Group Dynamics (Gruppe, Items 11–13): Measures peer collaboration, active collective engagement, and group cohesion. In command post operations, no individual acts in isolation; performance hinges on the fluid integration of functional units (S1 Personnel, S2 Intelligence, S3 Operations, S4 Logistics, S5 Media, S6 IT/Communications). Items evaluate active peer participation (Item 11), mutual operational support (Item 12), and perceived group cohesion (Item 13).
2. Learning Outcome Constructs (Ergebnisebene)
The outcome tier captures the subjective cognitive and operational transformations experienced by participants upon conclusion of the training:
- Competence Acquisition (Kompetenzerwerb, Items 14–19): Reflects self-perceived increases in functional and non-technical command competencies. Rooted in research on team learning and communication (Van den Bossche et al., 2006), this dimension focuses on cross-functional communication efficacy (Item 14), high-stress decision-making fluidity (Item 15), metacognitive awareness of personal psychological boundaries (Item 16), stress resilience and emotional calm (Item 17), peer information processing (Item 18), and the capacity to critically appraise intelligence received from colleagues (Item 19). These facets constitute the core socio-cognitive repertoire required in high-tempo crisis environments.
- Training Transfer (Transfer, Items 20–22): Evaluates perceived operational readiness and the generalizability of acquired competencies to real-world command duty. It operationalizes anticipatory operational readiness (Item 20), subjective psychological security gained through simulation exercises (Item 21), and the direct transferability of training content into future command staff deployments (Item 22). Transfer measures bridge institutional simulation with subsequent field survival and decision efficacy.
Theoretical Framework
The architectural blueprint of the FIRE-CU is anchored in three theoretical paradigms: Donald Kirkpatrick’s hierarchical evaluation taxonomy, Blanchard and Thacker’s dual-level training framework, and the theory of Shared Mental Models within High-Reliability Teams.
Kirkpatrick’s Four-Level Training Evaluation Model
The standard paradigm for educational evaluation in organizational settings is the four-level model developed by Donald Kirkpatrick (1979; Kirkpatrick & Kirkpatrick, 2006). This model posits that training efficacy must be systematically analyzed across a hierarchical continuum:
- Level 1: Reaction — Measures participants’ subjective impressions, affective responses, and cognitive valuation of the instructional event.
- Level 2: Learning — Assesses the acquisition of knowledge, procedural skills, attitudes, and operational competencies during the learning intervention.
- Level 3: Behavior — Gauges the extent to which on-the-job execution changes as a direct consequence of training (behavioral application and transfer).
- Level 4: Results — Evaluates organizational-level outcomes, systemic impact, safety metrics, and cost reductions.
Meta-analytic research by Alliger, Tannenbaum, Bennett, Traver, and Shotland (1997) refined Level 1 by differentiating between affective reactions (whether trainees enjoyed the course) and utility reactions (whether trainees perceived the training as instrumentally useful for their job performance). Alliger and colleagues demonstrated that utility reactions demonstrate significantly stronger predictive relationships with downstream learning (Level 2) and behavioral transfer (Level 3) than affective reactions. The FIRE-CU heavily prioritizes utility reactions within its process scales and captures immediate subjective capability growth corresponding to Kirkpatrick’s Level 2 and anticipatory Level 3 readiness.
Blanchard & Thacker’s Process-Outcome Dichotomy
To overcome the limitations of classic end-of-course satisfaction metrics, Blanchard and Thacker (2013) conceptualized training assessment through a functional bifurcation between Process Data and Outcome Data. Process evaluations monitor instructional delivery fidelity, trainer pedagogical conduct, structural consistency, and learning friction, diagnosing *why* a course succeeded or failed. In contrast, outcome evaluations verify whether pedagogical objectives were successfully achieved. Blanchard and Thacker argued that relying solely on outcome parameters conceals instructional mechanisms, while analyzing process parameters without outcome verification leaves learning gains unmeasured. The FIRE-CU explicitly operationalizes this model by treating Instructor Behavior, Requirement Level, Structure, and Group as process inputs, which directly facilitate or constrain the outcomes of Competence Acquisition and Transfer.
Shared Mental Models in High-Reliability Teams
At the command staff level, catastrophic operational failures rarely stem from technical incompetence alone; rather, they emerge from breakdowns in Shared Mental Models (SMM), distributed situation awareness, and inter-individual information processing (Cannon-Bowers et al., 1993; Salas et al., 2001). Members of a crisis staff must construct convergent cognitive frameworks regarding incident dynamics, tactical priorities, and task interdependencies under severe time compression and high ambiguity. Drawing from the work of Van den Bossche, Gijselaers, Segers, and Kirschner (2006) on team learning mechanisms, the FIRE-CU integrates specific items into its Competence Acquisition scale (Items 14, 18, and 19) that assess collaborative information exchange, mutual cognitive alignment, and collective sensemaking.
Validity
Empirical validation of the FIRE-CU was established during an extensive multi-cohort field investigation at the Institut der Feuerwehr NRW, encompassing instructional iterations over a 16-week cycle (Thielsch & Hadzihalilovic, 2020). The validation protocol verified content, construct, convergent, discriminant, and criterion-related validity.
Content Validity
Content validity was established through a multi-stage qualitative and quantitative development pipeline. Initial adaptations of the base FIRE instrument (Schulte & Thielsch, 2019) were subjected to an online pilot evaluation ($n = 8$) among active command course trainees. The results were reviewed in structured cognitive interviews with the Division Head and Deputy Division Head of the Crisis Management and Research Department at IdF NRW. Based on panel feedback, items reflecting non-applicable low-level tactical duties were discarded, three new competence items were integrated from validated team-learning inventories (Van den Bossche et al., 2006), and transfer items were adapted to command post simulation terminology. Items were adjusted to ensure gender-neutral German phrasing and cross-organizational applicability.
Construct, Convergent, and Discriminant Validity
Construct validity is substantiated by the confirmation of the hypothesized multi-tier latent dimensional structure via structural equation modeling. Confirmatory factor analyses demonstrated that the latent dimensions operate as conceptually coherent, empirically separable entities. The correlations among the process factors (Instructor Behavior, Structure, and Group Dynamics) were positive and statistically significant, while correlating negatively with Requirement Level (overload). This pattern supports theoretical expectations: elevated instructional clarity and pedagogical excellence attenuate perceived cognitive overload.
Convergent validity of the outcome scales is supported by substantial item-factor loadings ($p < .001$) across the Competence Acquisition and Transfer dimensions. Discriminant validity is evidenced by the distinct empirical separation between Process and Outcome latent variables. Process variables (e.g., Structure and Instructor Behavior) correlate moderately with Competence Acquisition and Transfer ($r \approx .35$ to $.55$), confirming that while instructional process quality fosters learning outcomes, it does not explain excessive shared variance that would indicate construct redundancy.
Criterion-Related Validity
Criterion validity was corroborated by correlating the six core psychometric subscales with external validation metrics, including the optional Global Assessment indicators (Item 23: overall learning gain; Item 24: academic school grade; Item 25: recommendation likelihood) and standardized post-training mood assessments. Trainees who reported higher instructor effectiveness, superior structural transparency, and higher group cohesion assigned significantly better overall school grades ($p < .001$) and demonstrated higher willingness to recommend the program to external emergency service peers. In parallel, elevated Competence Acquisition and Transfer scores predicted heightened post-exercise self-efficacy and positive affective states upon course completion.
Reliability
The psychometric reliability of the FIRE-CU was evaluated through comprehensive internal consistency analyses utilizing both classic coefficients (Cronbach’s $\alpha$) and composite reliability estimators (McDonald’s $\omega$), which accommodate violations of tau-equivalence in structural equation models.
Across the validation studies conducted at IdF NRW, the internal consistency values for all six core subscales demonstrated acceptable to excellent reliability, well exceeding standard psychometric thresholds for applied research ($\alpha > .70$):
- Instructor Behavior (Dozentenverhalten, 4 items): Cronbach’s $\alpha = .84$ to $.89$; McDonald’s $\omega = .85$ to $.90$. Demonstrates robust measurement precision regarding instructional quality.
- Requirement Level (Anforderungsniveau, 3 items): Cronbach’s $\alpha = .75$ to $.79$; McDonald’s $\omega = .76$ to $.80$. Reflects adequate homogeneity despite evaluating three distinct facets of overload (volume, tempo, difficulty).
- Structure (Struktur, 3 items): Cronbach’s $\alpha = .81$ to $.86$; McDonald’s $\omega = .82$ to $.87$. Demonstrates strong item convergence regarding organizational clarity.
- Group Dynamics (Gruppe, 3 items): Cronbach’s $\alpha = .76$ to $.82$; McDonald’s $\omega = .77$ to $.83$. Successfully measures peer cohesion and collaborative engagement.
- Competence Acquisition (Kompetenzerwerb, 6 items): Cronbach’s $\alpha = .86$ to $.90$; McDonald’s $\omega = .87$ to $.91$. Indicates exceptional measurement reliability across non-technical command staff competencies.
- Transfer (Transfer, 3 items): Cronbach’s $\alpha = .82$ to $.87$; McDonald’s $\omega = .83$ to $.88$. Captures perceived future field readiness with minimal measurement error.
Overall, across the full battery of subscales, Cronbach’s alpha spans $.75$ to $.90$ and McDonald’s omega ranges from $.76$ to $.91$. Test-retest reliability evaluations are theoretically contraindicated for immediate post-training satisfaction metrics, as repeating the survey at a later time introduces variance from real-world field exposures. However, the stability of the factor saturation and parallel internal consistency across independent training cohorts confirms the empirical stability of the measurement tool.
Factor Analysis
To substantiate the structural architecture of the FIRE-CU, confirmatory factor analyses (CFA) were conducted within R utilizing the structural equation modeling package lavaan (Rosseel, 2012), supplemented by packages psych (Revelle, 2018) and mice (van Buuren & Groothuis-Oudshoorn, 2011). Due to the theoretical bifurcation between instructional mechanisms and educational outcomes, two independent confirmatory measurement models were evaluated against empirical criteria (Hu & Bentler, 1999; Schermelleh-Engel et al., 2003).
1. Confirmatory Factor Analysis: Process Level (Prozessebene)
A four-factor oblique measurement model was specified for the 13 process items, allocating them to Instructor Behavior (4 items), Requirement Level (3 items), Structure (3 items), and Group (3 items). The baseline structural model yielded a commendable fit to the empirical data:
- $\chi^2(59) = 111.55, p < .001$
- Root Mean Square Error of Approximation ($\text{RMSEA}$) = $0.06$ ($90% \text{ CI } [0.04, 0.07]$)
- Standardized Root Mean Square Residual ($\text{SRMR}$) = $0.05$
- Comparative Fit Index ($\text{CFI}$) = $0.96$
- Tucker-Lewis Index ($\text{TLI}$) = $0.95$
While the overall $\chi^2$ statistic was significant (a known artifact of moderate-to-large sample sizes), the normed chi-square ratio was within the desirable range ($\chi^2 / \text{df} = 1.89$). Diagnostic inspection of modification indices suggested an empirical error covariance between Item 12 (“The participants supported each other”) and Item 13 (“I think there was good cohesion in the course”). Given their shared semantic focus on peer solidarity within the Group subscale, allowing this residual covariance was theoretically justified. The revised model demonstrated an enhanced fit:
- $\chi^2(58) = 96.45, p < .01$ ($\chi^2 / \text{df} = 1.66$)
- $\text{RMSEA} = 0.05$ ($90% \text{ CI } [0.03, 0.06]$)
- $\text{SRMR} = 0.04$
- $\text{CFI} = 0.97$
- $\text{TLI} = 0.96$
Standardized factor loadings across all 13 process items were uniformly high and statistically significant ($p < .001$), confirming solid convergent measurement validity for the instructional environment.
2. Confirmatory Factor Analysis: Outcome Level (Ergebnisebene)
The hypothesized two-factor oblique measurement model for the 9 outcome items (Competence Acquisition, 6 items; Transfer, 3 items) initially exhibited suboptimal fit indices under baseline constraints ($\chi^2(26) = 140.96, p < .001$; $\text{RMSEA} = 0.12$ [$0.11, 0.14$]; $\text{SRMR} = 0.07$; $\text{CFI} = 0.87$; $\text{TLI} = 0.82$).
Inspection of modification indices revealed specific inter-item residual covariances, primarily between Transfer Item 20 (“feel well prepared”) and Item 21 (“gained necessary confidence”), and among specific non-technical competence facets within the 6-item Competence Acquisition scale (correlations between Item 16 [boundary awareness] with Items 15, 18, and 19; and between Item 18 [information processing] and Item 19 [critical verification]). Theoretically, non-technical command staff competencies are systemic and interdependent (Lamers, 2016); processing incoming tactical information inevitably correlates with the capacity to critically assess that data. Incorporating these localized error covariances produced an exceptional model fit:
- $\chi^2(21) = 45.90, p < .01$ ($\chi^2 / \text{df} = 2.18$)
- $\text{RMSEA} = 0.06$ ($90% \text{ CI } [0.04, 0.09]$)
- $\text{SRMR} = 0.04$
- $\text{CFI} = 0.97$
- $\text{TLI} = 0.95$
These confirmatory investigations establish that the FIRE-CU operates as a psychometrically sound, multi-dimensional assessment instrument capturing empirical variance across both training delivery and operational skill mastery.
Instrument / Measurement Tool
The operational features, technical parameters, and diagnostic deployment procedures of the FIRE-CU are structured as follows:
- Tool Type: Standardized diagnostic training evaluation questionnaire / Post-training reaction and learning assessment tool.
- Target Population: Officers, command staff personnel, senior rescue team leaders, disaster management officials, and crisis cell coordinators within fire services, humanitarian relief organizations (THW, DRK), law enforcement, municipal crisis units, and military command bodies.
- Administration Mode: Standardized paper-and-pencil questionnaire or secure digital/mobile survey interface. Paper-and-pencil delivery is recommended in training centers due to higher response rates.
- Item Inventory:
- Core Psychometric Items: 22 standardized rating statements measuring the six primary dimensions.
- Supplemental Global Indicators (Optional): 3 single-item summary questions (Item 23: Subjective learning gain; Item 24: Traditional school grade [1–6]; Item 25: Recommendation intention [Yes/No]).
- Supplemental Qualitative Item (Optional): 1 open-ended feedback entry box (Item 26: Constructive remarks, praise, or structural criticism for instructors).
- Completion Duration: Approximately 4 to 5 minutes, minimizing cognitive fatigue at the conclusion of training.
- Authentic Response Scale:
Items 1 through 23 utilize an unforced, 7-point Likert agreement scale with an additional non-substantive alternative:
1= stimme gar nicht zu (Strongly disagree)2= stimme nicht zu (Disagree)3= stimme eher nicht zu (Somewhat disagree)4= neutral (Neutral)5= stimme eher zu (Somewhat agree)6= stimme zu (Agree)7= stimme vollkommen zu (Strongly agree)[ ]= nicht sinnvoll beantwortbar (Cannot be meaningfully answered / Not applicable)
Item 24 utilizes a 6-point German grading scale (1 = sehr gut / Very good, 2 = gut / Good, 3 = befriedigend / Satisfactory, 4 = ausreichend / Sufficient, 5 = mangelhaft / Deficient, 6 = ungenügend / Failing). Item 25 uses a binary format (ja / Yes; nein / No). Item 26 provides an open narrative text box.
- Scoring and Computational Rules:
- Item-Level Scoring: Numeric values from 1 to 7 are assigned directly to the Likert categories. Responses flagged as nicht sinnvoll beantwortbar receive no score and are treated as missing values.
- Requirement Level Inversion: For five of the scales (Instructor Behavior, Structure, Group, Competence, Transfer), higher scores indicate superior training quality. On the Requirement Level scale (Items 5–7), higher raw scores indicate trainee overload and excessive pacing. For organizational feedback reports, scores on Items 5–7 can be reverse-coded ($\text{Score}_{\text{recoded}} = 8 – \text{Score}_{\text{raw}}$) so that higher values consistently indicate better instructional quality across all scales.
- Scale Score Calculation: The scale score is calculated as the unweighted arithmetic mean of the valid items within that subscale:
- Instructor Behavior: $\text{Mean}(\text{Item } 1, 2, 3, 4)$
- Requirement Level: $\text{Mean}(\text{Item } 5, 6, 7)$
- Structure: $\text{Mean}(\text{Item } 8, 9, 10)$
- Group Dynamics: $\text{Mean}(\text{Item } 11, 12, 13)$
- Competence Acquisition: $\text{Mean}(\text{Item } 14, 15, 16, 17, 18, 19)$
- Transfer: $\text{Mean}(\text{Item } 20, 21, 22)$
- Handling of Global Items: Items 23 to 25 must never be combined into a composite score. Item 24 (School Grade) should be analyzed via median or frequency distributions due to its ordinal measurement level.
- Data Quality and Anonymity Threshold:
- To safeguard participant confidentiality in small command staff cohorts, aggregated group reports must only be generated when an anonymity threshold of at least $N = 8$ completed questionnaires is reached (or a minimum response rate of $50%$ in small seminars of 10 to 15 participants; Thielsch & Weltzin, 2013).
- Individual questionnaires with three or more missing core items should be excluded from final psychometric analyses.
Permissions & Fee and Test Year
The FIRE-CU was finalized and published in 2020 by Meinald T. Thielsch and Edin Hadzihalilovic in collaboration with the Institut der Feuerwehr NRW. The instrument is made available as an open-access, non-commercial psychological assessment tool for research, educational quality assurance, and organizational benchmarking within public safety entities and academic institutions.
Licensing and Usage Policy: Academic institutions, civil protection departments, fire service colleges, and non-profit emergency management organizations are permitted to utilize the FIRE-CU without royalty fees, provided appropriate scholarly attribution is maintained. Commercial distribution, inclusion in proprietary digital human resource suites, or fee-based training consulting using the instrument requires formal written authorization from the primary authors and the University of Münster (OWMS). Comprehensive test documentation, German and English questionnaire forms, and scoring templates can be accessed via the official project repository at the University of Münster Project FIRE Portal.
References
- Alliger, G. M., Tannenbaum, S. I., Bennett, W., Traver, H., & Shotland, A. (1997). A meta-analysis of the relations among training criteria. Personnel Psychology, 50(2), 341–358. https://doi.org/10.1111/j.1744-6570.1997.tb00911.x
- Arthur, W., Bennett, W., Edens, P. S., & Bell, S. T. (2003). Effectiveness of training in organizations: A meta-analysis of design and evaluation features. Journal of Applied Psychology, 88(2), 234–245. https://doi.org/10.1037/0021-9010.88.2.234
- Blanchard, P. N., & Thacker, J. W. (2013). Effective training: Systems, strategies, and practices (5th ed.). Pearson.
- Cannon-Bowers, J. A., Salas, E., & Converse, S. (1993). Shared mental models in expert team decision making. In N. J. Castellan, Jr. (Ed.), Individual and group decision making: Current issues (pp. 221–246). Lawrence Erlbaum Associates.
- Gediga, G., Hamborg, K.-C., & Willumeit, K. (2000). Das Kieler Evaluationsinstrument für Lehrveranstaltungen (KIEL). Universität Osnabrück.
- Grohmann, A., & Kauffeld, S. (2013). Evaluating training outcomes: A validate measure to assess training transfer. International Journal of Training and Development, 17(2), 135–151. https://doi.org/10.1111/ijtd.12005
- Hagemann, V., Kluge, A., & Ritzmann, S. (2012). Teamressourcen im Notfall: Training und Evaluation von proaktiven Verhaltensweisen zur Erhöhung der Patientensicherheit. Zeitschrift für Arbeits- und Organisationspsychologie, 56(3), 119–137. https://doi.org/10.1026/0932-4089/a000084
- Heath, R. (1998). Crisis management for managers and executives. Financial Times Pitman.
- Heimann, R., & Hofinger, G. (2016). Stäbe und Stabsarbeit. In G. Hofinger & R. Heimann (Eds.), Handbuch Stabsarbeit: Führungs- und Krisenstäbe in Einsatz und Verwaltung (pp. 3–14). Springer. https://doi.org/10.1007/978-3-662-48187-5_1
- Hu, L. T., & Bentler, P. M. (1999). Cutoff criteria for fit indexes in covariance structure analysis: Conventional criteria versus new alternatives. Structural Equation Modeling: A Multidisciplinary Journal, 6(1), 1–55. https://doi.org/10.1080/10705519909540118
- Kirkpatrick, D. L. (1979). Techniques for evaluating training programs. Training and Development Journal, 33(6), 78–92.
- Kirkpatrick, D. L., & Kirkpatrick, J. D. (2006). Evaluating training programs: The four levels (3rd ed.). Berrett-Koehler.
- Lamers, T. (2016). Besonderheiten der Stabsarbeit bei der Feuerwehr. In G. Hofinger & R. Heimann (Eds.), Handbuch Stabsarbeit: Führungs- und Krisenstäbe in Einsatz und Verwaltung (pp. 31–42). Springer. https://doi.org/10.1007/978-3-662-48187-5_3
- Passmore, J., & Joao Velez, M. (2014). Training evaluation. In K. Kraiger, J. Passmore, N. R. dos Santos, & S. Malvezzi (Eds.), The Wiley Blackwell handbook of the psychology of training, development, and performance improvement (pp. 136–153). Wiley-Blackwell. https://doi.org/10.1002/9781118736982.ch8
- Queck, C., & Gonner, G. (2016). Führung und Zusammenarbeit im Führungsstab der Feuerwehr. In G. Hofinger & R. Heimann (Eds.), Handbuch Stabsarbeit: Führungs- und Krisenstäbe in Einsatz und Verwaltung (pp. 143–154). Springer. https://doi.org/10.1007/978-3-662-48187-5_11
- Revelle, W. (2018). psych: Procedures for personality and psychological research (R package version 1.8.12). Northwestern University. https://cran.r-project.org/package=psych
- Röseler, S., Thielsch, M. T., & Schulte, N. P. (2020). Feedback zur Rettungskräfteentwicklung – Einsatzübung (FIRE-E). Universität Münster. https://doi.org/10.17605/OSF.IO/7VY2D
- Rosseel, Y. (2012). lavaan: An R package for structural equation modeling. Journal of Statistical Software, 48(2), 1–36. https://doi.org/10.18637/jss.v048.i02
- Salas, E., & Cannon-Bowers, J. A. (2001). The science of training: A decade of progress. Annual Review of Psychology, 52(1), 471–499. https://doi.org/10.1146/annurev.psych.52.1.471
- Schermelleh-Engel, K., Moosbrugger, H., & Müller, H. (2003). Evaluating the fit of structural equation models: Tests of significance and descriptive goodness-of-fit measures. Methods of Psychological Research Online, 8(2), 23–74.
- Schulte, N. P., & Thielsch, M. T. (2019). Evaluation der Führungsausbildung bei der Feuerwehr: Entwicklung und Validierung des Fragebogens FIRE. Zeitschrift für Arbeits- und Organisationspsychologie, 63(4), 195–209. https://doi.org/10.1026/0932-4089/a000307
- Staufenbiel, T. (2000). Fragebogen zur Evaluation von universitären Lehrveranstaltungen (FEVOR). Diagnostica, 46(4), 169–181. https://doi.org/10.1026//0012-1924.46.4.169
- Thielsch, M. T., & Hadzihalilovic, E. (2020). Feedback zur Rettungskräfteentwicklung – Führungsstab (FIRE-CU). Universität Münster. https://doi.org/10.17605/OSF.IO/UFXNE
- Thielsch, M. T., & Hirschfeld, G. (2012). Münsteraner Fragebogen zur Evaluation von Lehrveranstaltungen im Schüler-Feedback (MFE-Sr). Empirische Pädagogik, 26(4), 450–467.
- Thielsch, M. T., & Weltzin, S. (2013). Evaluation von eLearning: Ein Leitfaden zur Praxis. Waxmann.
- van Buuren, S., & Groothuis-Oudshoorn, K. (2011). mice: Multivariate imputation by chained equations in R. Journal of Statistical Software, 45(3), 1–67. https://doi.org/10.18637/jss.v045.i03
- Van den Bossche, P., Gijselaers, W. H., Segers, M., & Kirschner, P. A. (2006). Social and cognitive factors driving teamwork in collaborative learning environments: Team learning beliefs and behaviors. Small Group Research, 37(5), 490–521. https://doi.org/10.1177/1046496406292938
- Wybo, J. L., & Kowalski, K. M. (1998). Command centers and emergency management: An overview. Disaster Prevention and Management: An International Journal, 7(2), 110–116. https://doi.org/10.1108/09653569810216091