Abstract
The Clinical Outcomes in Routine Evaluation—Outcome Measure (CORE-OM) is a standardized, pan-theoretical, self-report psychometric instrument designed to assess psychological distress and therapeutic change in adult populations undergoing psychotherapy, counseling, and allied mental health interventions. Developed in the United Kingdom by Evans and colleagues (2000, 2002) under the auspices of the CORE System Group, the CORE-OM addresses the long-standing clinical and empirical necessity for a universal, practitioner-friendly measurement tool that bridges the gap between routine clinical audit, randomized controlled trials, and practice-based research networks. The instrument comprises 34 items evaluated on a 5-point Likert scale (ranging from 0 = “Not at all” to 4 = “Most or all the time”) referencing the respondent’s experiences over the preceding week. The measure spans four primary conceptual domains: Subjective Well-being (4 items), Problems/Symptoms (12 items covering anxiety, depression, physical symptoms, and trauma), Life Functioning (12 items evaluating general functioning, close personal relationships, and social interactions), and Risk/Harm (6 items assessing risk to self and risk to others). Across extensive normative and clinical psychometric evaluations, the CORE-OM has demonstrated robust internal consistency (Cronbach’s alpha ranging from .75 to .94 across subscales and exceeding .94 for the total score), high test-retest reliability ($r = .87$ to $.91$), and exemplary convergent validity with established single-construct instruments such as the Beck Depression Inventory (BDI) and the Symptom Checklist-90-Revised (SCL-90-R). Its factor structure yields an overarching general psychological distress dimension alongside discernible sub-dimensions, providing clinicians and researchers with both macro-level tracking of treatment efficacy and granular profiles of specific psychological difficulties.
Keywords
CORE-OM, Clinical Outcomes in Routine Evaluation, outcome measurement, psychotherapy outcome, psychological distress, routine outcome monitoring, subjective well-being, life functioning, risk assessment, psychometrics
Authors
The CORE-OM was developed through a multi-institutional collaborative consortium in the United Kingdom known as the CORE System Group. The principal architects of the instrument include:
- Chris Evans, MRCPsych, MSc — Consultant Psychiatrist in Psychotherapy, Institute of Group Analysis; Nottinghamshire Healthcare NHS Trust; Mandala Centre, Nottingham, UK. Contact: [email protected].
- John Mellor-Clark, PhD — Research Fellow and Director of CORE IMS, Psychological Therapies Research Centre, University of Leeds, UK.
- Frank Margison, FRCPsych — Consultant Psychiatrist in Psychotherapy, Manchester Mental Health and Social Care Trust; University of Manchester, UK.
- Michael Barkham, PhD — Professor of Clinical Psychology, Department of Psychology, University of Sheffield; formerly Director of the Psychological Therapies Research Centre, University of Leeds, UK.
- Janice Connell, MSc — Research Fellow, Psychological Therapies Research Centre, University of Leeds, UK.
- Kerry Audin, PhD — Research Fellow, Psychological Therapies Research Centre, University of Leeds, UK.
- Graeme McGrath, FRCPsych — Consultant Psychiatrist, Manchester Mental Health Partnership, Manchester, UK.
Purpose
The primary purpose of the Clinical Outcomes in Routine Evaluation—Outcome Measure is to provide a generic, broad-spectrum assessment tool for routine administration in psychological therapy settings across national health services, university counseling services, primary care mental health networks, and independent practices. Prior to the development of the CORE system in the late 1990s, psychotherapy research and quality assurance suffered from extreme fragmentation. Clinical trials relied predominantly on proprietary, disorder-specific assessment batteries (such as specific inventories for major depression or panic disorder) that failed to capture comorbid difficulties, life functioning, or existential well-being, while routine clinical audit initiatives employed idiosyncratic, non-standardized feedback forms that precluded cross-service comparisons.
The CORE-OM was designed with three foundational objectives:
- Routine Outcome Monitoring (ROM): Enabling therapists and service managers to track individual patient trajectories session-by-session or pre-to-post treatment. By comparing raw change against statistically derived thresholds for Reliable and Clinically Significant Change (Jacobson & Truax, 1991), practitioners can identify treatment failure, sudden gains, or clinical deterioration in real time.
- Clinical Governance and Service Evaluation: Furnishing mental health services with benchmarked, aggregated data capable of demonstrating accountability, equity of access, and effectiveness across diverse treatment modalities, including psychodynamic, cognitive-behavioral, humanistic, and systemic therapies.
- Clinical Triage and Risk Screening: Providing explicit, immediate alerts regarding patient safety via dedicated items assessing suicidal ideation, suicide plans, self-harm, and interpersonal violence or intimidation.
From a theoretical and methodological perspective, the CORE-OM bridges the gap between academic efficacy trials (investigating whether an intervention works under ideal, experimental parameters) and real-world effectiveness research (determining how diverse client populations respond to psychological interventions in naturalistic conditions). Its universal applicability across diagnoses makes it suitable for heterogeneous clinical presentations where comorbidity and diffuse distress predominate.
Psychological Construct
The CORE-OM operationalizes psychological distress and functioning through a multidimensional architecture comprising four primary conceptual domains and several nested sub-facets:
1. Subjective Well-Being (4 Items)
This domain captures the client’s core affective state, optimism, and fundamental sense of self-worth. In contrast to purely symptom-focused pathology models, well-being reflects the positive psychological resources and baseline subjective state essential for recovery. It includes assessments of self-esteem (e.g., item 4: “I have felt O.K. about myself”), future orientation and hopefulness (e.g., item 31: “I have felt optimistic about my future”), and general affective equilibrium.
2. Problems / Symptoms (12 Items)
The problem domain evaluates internal psychological distress across four key clinical clusters reflecting common mental health complaints:
- Depression / Low Mood (4 items): Reflecting pervasive sadness, crying spells, exhaustion, and despair (e.g., item 5: “I have felt totally lacking in energy and enthusiasm”; item 14: “I have felt like crying”; item 23: “I have felt despairing or hopeless”).
- Anxiety (4 items): Encompassing somatic tension, psychological apprehension, and catastrophic physiological arousal (e.g., item 2: “I have felt tense, anxious or nervous”; item 15: “I have felt panic or terror”).
- Physical / Somatic Problems (2 items): Capturing sleep disturbance and pain or bodily discomfort not solely attributed to organic disease (e.g., item 8: “I have been troubled by aches, pains or other physical problems”; item 18: “I have had difficulty getting to sleep or staying asleep”).
- Trauma / Cognitive Intrusion (2 items): Probing intrusive memories, flashbacks, and cognitive disruptions characteristic of traumatic stress (e.g., item 13: “I have been disturbed by unwanted thoughts and feelings”; item 28: “Unwanted images or memories have been distressing me”).
3. Life Functioning (12 Items)
This dimension examines the pragmatic, interpersonal, and social expressions of psychological distress. Drawing on systemic and social role theories, functioning is evaluated across three interactive spheres:
- General / Daily Role Functioning (4 items): The capacity to execute occupational, domestic, or academic responsibilities and manage problems (e.g., item 21: “I have been able to do most things I needed to”; item 20: “My problems have been impossible to put to one side”).
- Close Personal Relationships (4 items): The client’s perceived social support, intimacy, and feelings of interpersonal isolation (e.g., item 1: “I have felt terribly alone and isolated”; item 19: “I have felt warmth or affection for someone”).
- Social and Community Functioning (4 items): The ability to engage with others, tolerate social scrutiny, and navigate community life without debilitating shame or irritability (e.g., item 10: “Talking to people has felt too much for me”; item 29: “I have been irritable when with other people”).
4. Risk / Harm (6 Items)
The risk domain acts as a critical clinical safeguard, evaluating behaviors and ideations that pose imminent danger to life or physical well-being. Unlike the other domains, risk items tend to exhibit severe positive skew in general populations but carry high clinical urgency. The domain bifurcates into:
- Risk to Self (4 items): Covering passive suicidal ideation, explicit self-harm thoughts, active plans, and hazardous behaviors (e.g., item 9: “I have thought of hurting myself”; item 16: “I made plans to end my life”; item 24: “I have thought it would be better if I were dead”; item 34: “I have hurt myself physically or taken dangerous risks with my health”).
- Risk to Others (2 items): Assessing overt interpersonal violence and intimidation (e.g., item 6: “I have been physically violent to others”; item 22: “I have threatened or intimidated another person”).
Theoretical Framework
The conceptual foundation of the CORE-OM is anchored in contemporary clinical psychology paradigms, particularly Kenneth Howard’s Phase Model of Psychotherapy Outcome (Howard et al., 1993) and the broader framework of Practice-Based Evidence (PBE) championed by Barkham and colleagues (Barkham et al., 2010).
Howard’s Phase Model of Psychotherapeutic Change
Howard and colleagues proposed that psychotherapeutic recovery does not occur as an undifferentiated, uniform reduction in pathology, but rather progresses through three sequential, logically ordered phases:
- Remoralization: In the initial sessions of therapy, the primary mechanism of change is the enhancement of subjective well-being and the restoration of hope. Despair is alleviated as the patient establishes therapeutic rapport. The CORE-OM Subjective Well-being domain directly maps onto this remoralization phase.
- Remediation: As therapy deepens, the focus shifts toward symptom reduction and the resolution of specific clinical complaints (such as panic, insomnia, dysphoria, and traumatic flashbacks). This corresponds directly to the CORE-OM Problems/Symptoms domain.
- Rehabilitation: The final and most prolonged stage involves unlearning maladaptive relational patterns, establishing social functioning, and restoring interpersonal effectiveness in community and vocational domains. This phase is tracked precisely by the CORE-OM Life Functioning domain.
By encompassing all three phases of Howard’s model, the CORE-OM provides an empirical architecture capable of registering treatment progress regardless of whether a patient is in brief crisis intervention (primarily targeting remoralization and risk containment) or long-term dynamic or cognitive reconstruction (targeting deep remediation and rehabilitation).
Practice-Based Evidence vs. Evidence-Based Practice
Historically, the empirical validation of psychological therapies relied almost exclusively on Evidence-Based Practice (EBP) derived from tightly controlled, randomized clinical trials (RCTs). While RCTs offer high internal validity, they often exclude complex patients with severe comorbidity, chronic relational pathology, or high self-harm risk. The CORE System was constructed to foster Practice-Based Evidence, where standardized measurement is integrated directly into routine, non-selective clinical care. In this paradigm, the instrument must simultaneously function as a reliable psychometric scale for researchers and an unobtrusive, clinically meaningful conversation starter for practitioners.
Validity
The psychometric validity of the CORE-OM has been thoroughly established through rigorous investigations involving tens of thousands of participants across clinical and non-clinical cohorts (Evans et al., 2002; Connell et al., 2007; Lyne et al., 2006).
Construct and Convergent Validity
Convergent validity has been evaluated against gold-standard psychiatric and psychological scales. In early validation studies by Evans et al. (2002), the CORE-OM Total Score exhibited strong correlations with:
- Beck Depression Inventory (BDI): $r = .73$ to $.81$, indicating that depressive distress is prominently represented within the general problem dimension.
- Symptom Checklist-90-Revised (SCL-90-R) Global Severity Index (GSI): $r = .77$ to $.88$, establishing parity with comprehensive psychiatric symptom inventories.
- Clinical Outcomes Assessment (OQ-45.2): $r = .85$, reflecting exceptional structural alignment between the two predominant international routine outcome systems.
- Inventory of Interpersonal Problems (IIP-32): Moderate-to-high correlations ($r = .58$ to $.65$) specifically with the CORE-OM Life Functioning and Interpersonal subscales.
Discriminant and Known-Groups Validity
The CORE-OM demonstrates exceptional known-groups validity, clearly differentiating between clinical populations seeking psychotherapy and non-clinical community or student populations. In standard validation cohorts (Evans et al., 2002):
- Clinical Population Mean: Pre-treatment clinical samples routinely score between $1.65$ and $1.90$ on the CORE-OM Total Score (standard deviation $\approx 0.65$).
- Non-Clinical General Population Mean: Community samples yield a mean Total Score between $0.50$ and $0.75$ (standard deviation $\approx 0.45$).
- Effect Size: The standardized mean difference (Cohen’s $d$) between clinical and non-clinical cohorts exceeds $1.8$, representing an extraordinarily robust discriminative boundary.
Sensitivity to Change and Criterion Validity
The instrument is highly sensitive to therapeutic change over time. Longitudinal studies demonstrate that successful completion of evidence-based psychological treatment yields large effect sizes (Cohen’s $d = 1.0$ to $1.4$) on the Total Score and Non-Risk Score. Furthermore, the instrument effectively differentiates between patients classified as recovered, improved, unchanged, or deteriorated according to independent therapist ratings and global change criteria.
Reliability
The reliability of the CORE-OM has been extensively demonstrated across numerous large-scale clinical trials, university counseling surveys, and primary care cohorts.
Internal Consistency
Internal consistency estimates (Cronbach’s alpha) across standard validation studies consistently surpass conventional psychometric thresholds for individual clinical decision-making:
- Total Score (all 34 items): $\alpha = .94$ in clinical populations; $\alpha = .92$ to $.94$ in non-clinical cohorts.
- Non-Risk Score (items excluding risk, 28 items): $\alpha = .94$.
- Subjective Well-being Domain (4 items): $\alpha = .75$ to $.82$.
- Problems / Symptoms Domain (12 items): $\alpha = .88$ to $.90$.
- Life Functioning Domain (12 items): $\alpha = .86$ to $.88$.
- Risk / Harm Domain (6 items): $\alpha = .75$ to $.79$. (Note: While slightly lower than other domains, this is expected due to the low base rate and multi-directional nature of risk behaviors, which encompass both self-directed suicidal acts and outward aggression).
Test-Retest Reliability
In non-clinical samples assessed across a 1- to 2-week interval without psychological intervention, the CORE-OM demonstrates high temporal stability:
- Total Score: Pearson correlation coefficient $r = .87$ to $.91$.
- Domain Subscales: Test-retest correlations range from $r = .83$ (Risk) to $r = .89$ (Problems/Symptoms and Life Functioning).
Reliable Change Index (RCI)
Based on Jacobson and Truax’s classical psychometric methodology, the Reliable Change Index (RCI) value for the CORE-OM Total Score is calculated at approximately $0.5$ (on the 0–4 scale) or $5$ points (on a $0–40$ scaled score). Any pre-to-post change exceeding this threshold can be classified with 95% statistical confidence as reflecting true therapeutic movement rather than measurement error.
Factor Analysis
The latent structural integrity of the CORE-OM has been the subject of extensive structural equation modeling (SEM), exploratory factor analysis (EFA), and confirmatory factor analysis (CFA) across diverse international populations.
Hierarchical and Bifactor Formulations
While the CORE-OM was designed with four conceptual domains (Well-being, Problems, Functioning, Risk), factor analytic studies indicate that empirical items from Well-being, Problems, and Functioning are heavily dominated by a single, overarching latent dimension of Global Psychological Distress (often referred to as a “super-factor” or general $g$-factor of distress). Exploratory factor analyses routinely reveal that the first unrotated factor accounts for over 35% to 45% of the total item variance.
Subsequent confirmatory factor analyses comparing unidimensional, multi-factor, and bifactor models (e.g., Lyne et al., 2006; Bedford et al., 2010) show that a bifactor model provides the best empirical fit to the data:
- Model Fit Indices: Comparative Fit Index (CFI) $> .92$, Tucker-Lewis Index (TLI) $> .91$, and Root Mean Square Error of Approximation (RMSEA) $< .05$.
- Factor Loadings: Non-risk items (1–5, 7–8, 10–15, 17–21, 23, 25–33) load strongly on the general distress factor (loadings typically ranging from $.55$ to $.82$).
- Independence of Risk: The Risk/Harm items load distinctly on a separate, secondary dimension. Self-harm/suicide items (items 9, 16, 24, 34) and outward violence items (items 6, 22) show lower loadings on the general distress factor, confirming that risk behavior forms an etiologically and empirically discrete phenomenon that must be scored and monitored independently of general psychological demoralization.
Instrument / Measurement Tool
The CORE-OM is a structured, pen-and-paper or digital psychological measurement tool possessing the following operational specifications:
- Test Type: Standardized self-report psychometric rating scale / routine outcome measure (ROM).
- Administration Mode: Self-administered (paper-and-pencil or online psychometric platforms); can also be clinician-administered or read aloud when cognitive or literacy limitations are present.
- Target Population: Adults (aged 18 to 65+). Parallel versions exist for young people (YP-CORE) and individuals with learning disabilities (LD-CORE).
- Completion Time: Approximately 5 to 10 minutes (34 items).
- Reference Window: The preceding week (“over the last week”).
- Response Scale: 5-point Likert scale:
- 0: Not at all
- 1: Only Occasionally
- 2: Sometimes
- 3: Often
- 4: Most or all the time
- Reverse-Scored Items (Positively Phrased): Items 3, 4, 7, 12, 19, 21, 31, and 32 are positively keyed and must be reverse-coded prior to subscale and total computation ($0 \rightarrow 4$, $1 \rightarrow 3$, $2 \rightarrow 2$, $3 \rightarrow 1$, $4 \rightarrow 0$).
- Scoring and Metrics:
- Mean Score (Item Average, Range 0–4): Calculated by dividing the sum of validly completed items by the total number of validly answered items. This is the metric most widely cited in clinical research.
- Clinical Score (Range 0–40): Calculated by multiplying the Mean Score by 10.
- Non-Risk Score (Items 1–5, 7–8, 10–15, 17–21, 23, 25–33): Represents the mean of the 28 non-risk items; highly recommended because the episodic nature of risk items can mask therapeutic gains in core functioning.
- Risk Score (Items 6, 9, 16, 22, 24, 34): Evaluated separately to track immediate safety and harm profiles.
- Handling Missing Data: If up to 10% of items are missing (e.g., $le 3$ items across the full scale), the mean is computed based on completed items. If more than 3 items are omitted across the 34 items, the total score should not be formally computed.
- Clinical Cut-Offs:
- General Population vs. Clinical Cut-off: Total mean score of $1.00$ (or Clinical Score of $10.0$). A client scoring $ge 1.00$ falls within the clinical range.
- Severity Bandings: Healthy / Non-clinical ($< 1.0$), Mild ($1.0 le text{score} < 1.5$), Moderate ($1.5 le text{score} < 2.0$), Moderate-to-Severe ($2.0 le text{score} < 2.5$), Severe ($ge 2.5$).
Permissions & Fee and Test Year
The CORE-OM was formally published in 2000 (initial validation) and 2002 (comprehensive psychometric establishment). It was deliberately conceived by its creators as an “open-access” public domain assessment tool to remove commercial barriers to scientific outcome evaluation.
- Copyright & Ownership: Copyright is held by the CORE System Trust (a registered charity in the United Kingdom).
- Licensing and Fees: The measure is free to reproduce on paper and use in clinical, research, and non-commercial educational settings without paying licensing fees, under a Creative Commons Attribution-NoDerivatives (CC BY-ND) framework.
- Integrity Stipulation: Users are strictly prohibited from changing item wording, altering the response scale, or deleting items if the instrument is to be called the “CORE-OM”. Commercial software developers incorporating the instrument into electronic health records (EHR) must seek approval from CORE System Trust / CORE Information Management Systems (CORE IMS).
- Official Contact and Distribution: Instruments, translations into over 30 languages, and scoring guidelines are accessible via the official website of the CORE System Trust (www.coresystemtrust.org.uk) and through Dr. Chris Evans ([email protected]).
References
The foundational scientific literature establishing the CORE-OM and its theoretical underpinnings includes:
- Barkham, M., Mellor-Clark, J., Connell, J., & Cahill, J. (2006). Making evidence-based practice practice-based: The CORE system, taxonomy of tools, and offering a collaborative agenda. Psychotherapy Research, 16(4), 443–456. https://doi.org/10.1080/10503300600650965
- Barkham, M., Hardy, G. E., & Mellor-Clark, J. (Eds.). (2010). Developing and delivering practice-based evidence: A guide for the psychological therapies. John Wiley & Sons. https://doi.org/10.1002/9780470688083
- Bedford, A., Watson, R., Evans, C., & Slack, D. (2010). The Clinical Outcomes in Routine Evaluation-Outcome Measure (CORE-OM): A dynamic polytomous Rasch model study. Journal of Applied Measurement, 11(2), 209–219.
- Connell, J., Barkham, M., Stiles, W. B., Twigg, E., Singleton, N., Evans, C., & Miles, J. N. (2007). Distribution of CORE-OM scores in a general population, clinical cut-off points and the effect of gender. The British Journal of Psychiatry, 190(4), 358–359. https://doi.org/10.1192/bjp.bp.105.017657
- Evans, C., Mellor-Clark, J., Margison, F., Barkham, M., Audin, K., Connell, J., & McGrath, G. (2000). CORE: Clinical Outcomes in Routine Evaluation. Journal of Mental Health, 9(3), 247–255. https://doi.org/10.1080/jmh.9.3.247.255
- Evans, C., Connell, J., Barkham, M., Margison, F., McGrath, G., Mellor-Clark, J., & Audin, K. (2002). Towards a standardised brief outcome measure: Psychometric properties and utility of the CORE-OM. The British Journal of Psychiatry, 180(1), 51–60. https://doi.org/10.1192/bjp.180.1.51
- Howard, K. I., Lueger, R. J., Maling, M. S., & Martinovich, Z. (1993). A phase model of psychotherapy outcome: Empirical support and tools for practice. Journal of Consulting and Clinical Psychology, 61(4), 678–685. https://doi.org/10.1037/0022-006X.61.4.678
- Jacobson, N. S., & Truax, P. (1991). Clinical significance: A statistical approach to defining meaningful change in psychotherapy research. Journal of Consulting and Clinical Psychology, 59(1), 12–19. https://doi.org/10.1037/0022-006X.59.1.12
- Lyne, K., Prowse, M., Swales, M., & Evans, C. (2006). The validity of the Clinical Outcomes in Routine Evaluation (CORE-OM) in a Welsh mental health setting. Clinical Psychology & Psychotherapy, 13(3), 201–209. https://doi.org/10.1002/cpp.491