Clinical PsychologyPsychiatric AssessmentPsychometrics

Patient Health Questionnaire – 9 (PHQ-9)

A comprehensive academic psychometric review of the Patient Health Questionnaire – 9 (PHQ-9), examining its theoretical foundations, diagnostic validity, factor structure, reliability, and clinical administration parameters.

memjavad
PUBLISHED
Scientifically Reviewed · Dr. Marwa Abd-Alazim · September 4, 2026
Medically & Scientifically Reviewed Verified: September 4, 2026
Dr. Marwa Abd-Alazim Ph.D.
Professor of Psychology University of Kerbala
Review Criteria & Clinical Standards

This content undergoes rigorous scientific peer-review and medical editorial standards at Arab Psychology Network to ensure clinical accuracy, validity, and compliance with evidence-based guidelines from leading psychological and healthcare authorities (APA / WHO).

1. Abstract

The Patient Health Questionnaire – 9 (PHQ-9) is a multipurpose, nine-item psychometric self-report instrument designed for screening, diagnosing, and measuring the symptom severity of major depressive disorder in clinical and epidemiological settings. Direct-mapped to the nine diagnostic criteria established in the Diagnostic and Statistical Manual of Mental Disorders, Fourth Edition (DSM-IV), the instrument assesses depressive symptomatology experienced over the preceding two-week reference window. Items assess anhedonia, depressed mood, sleep disturbance, fatigue, appetite changes, low self-worth, cognitive concentration difficulties, psychomotor agitation or retardation, and suicidal ideation. Each item is scored on a four-point frequency Likert scale ranging from 0 (“Not at all”) to 3 (“Nearly every day”), yielding a composite global severity score between 0 and 27. Clinically established thresholds categorize depression severity into minimal (0–4), mild (5–9), moderate (10–14), moderately severe (15–19), and severe (20–27), with a score of 10 or greater representing the standard cutoff for major depressive episodes.

Extensive psychometric investigations across diverse healthcare environments—including primary care clinics, psychiatric outpatient services, inpatient medical populations, and community cohorts—have documented exemplary measurement properties. The scale consistently demonstrates robust internal consistency (Cronbach’s alpha typically ranging from 0.86 to 0.89), strong test-retest reliability ($r = 0.84$ to $0.92$), high convergent validity with standardized psychiatric interviews (e.g., Structured Clinical Interview for DSM-IV), and strong predictive validity in tracking therapeutic trajectories. Confirmatory factor analyses generally substantiate a predominant unidimensional depression construct, with occasional secondary findings supporting a bifactor model delineating somatic and affective-cognitive sub-dimensions. The instrument’s brevity, clinical sensitivity, alignment with formal nosology, and public domain accessibility make it a cornerstone in global psychiatric assessment, clinical trials, and epidemiological research.

2. Keywords

Patient Health Questionnaire, PHQ-9, Major Depressive Disorder, Depression Screening, Psychometrics, DSM-IV Criteria, Severity Assessment, Primary Care Psychiatry, Clinical Measurement, Construct Validity

3. Authors

The Patient Health Questionnaire – 9 was developed by a team of prominent psychiatric and clinical epidemiological researchers:

  • Kurt Kroenke, M.D., MACP: Senior Research Scientist at the Regenstrief Institute, Professor of Medicine at Indiana University School of Medicine, Indianapolis, IN, USA. Dr. Kroenke is an internationally recognized expert in physical and psychological symptom management, measurement-based care, and the design of brief clinical assessment instruments.
  • Robert L. Spitzer, M.D.: Former Professor of Psychiatry at Columbia University and Chief of the Biometrics Research Department at the New York State Psychiatric Institute, New York, NY, USA. Dr. Spitzer was the lead architect of modern operationalized psychiatric nosology, serving as Chair of the American Psychiatric Association Task Force on DSM-III and DSM-III-R.
  • Janet B. W. Williams, Ph.D., D.S.W.: Professor Emerita of Clinical Psychiatric Social Work (in Psychiatry) at Columbia University College of Physicians and Surgeons, New York, NY, USA. Dr. Williams was an instrumental developer of diagnostic assessment tools including the Structured Clinical Interview for DSM (SCID) and the Prime-MD diagnostic series.

The developmental and validation field trials of the PHQ-9 were conducted in collaboration with research consortia across academic medical centers and funded in part through educational grants from Pfizer Inc., which subsequently placed the resulting PHQ instruments into the public domain without commercial restrictions.

4. Purpose

The fundamental clinical and scientific rationale underpinning the development of the PHQ-9 was the urgent necessity for a standardized, time-efficient, psychometrically rigorous, and clinically actionable depression measure suitable for high-volume primary care and specialized psychiatric environments. Epidemiological research conducted during the 1980s and 1990s revealed that primary care physicians underrecognized or misdiagnosed more than 50% of patients presenting with major depression, largely due to time constraints, the absence of systematic screening mechanisms, and the presence of confounding somatic complaints. Earlier diagnostic measures, such as the original clinician-administered PRIME-MD, improved detection rates but remained burdensome because they required extensive clinician interview time.

The PHQ-9 was specifically designed to overcome these operational limitations by functioning simultaneously as a dual-purpose instrument: a categorical diagnostic screening tool and a dimensional continuous severity metric. Categorically, the instrument’s nine items correspond 1:1 with the nine formal diagnostic criteria for Major Depressive Episode specified in DSM-IV and maintained in DSM-5. This direct nosological correspondence allows clinicians to determine whether a patient satisfies the syndromal criteria (requiring at least five of the nine symptoms present for “more than half the days” over the past two weeks, with at least one symptom being depressed mood or anhedonia). Dimensionally, by summing the score across all nine items, clinicians obtain a continuous metric of depression severity that enables precise staging, clinical risk stratification, and longitudinal tracking of treatment responsiveness.

In contemporary research paradigms, the PHQ-9 serves as a primary or secondary outcome variable across pharmacological, psychotherapeutic, digital health, and behavioral interventions. Its widespread implementation facilitates systematic benchmarking in comparative effectiveness studies and meta-analyses. Furthermore, the tenth non-scored functional impairment item (“How difficult have these problems made it for you to do your work, take care of things at home, or get along with other people?”) allows clinicians and investigators to gauge functional morbidity alongside subjective symptom severity, consistent with biopsychosocial clinical models.

5. Psychological Construct

The construct assessed by the PHQ-9 is operationalized as depressive episode symptomatology defined by contemporary psychiatric classification systems. Rather than viewing depression solely as an unobservable theoretical entity, the PHQ-9 operationalizes depression as a constellation of affective, cognitive, vegetative, and psychomotor disturbances. Each item is anchored to a distinct diagnostic criterion:

  • Anhedonia (Item 1): Evaluates markedly diminished interest or pleasure in all, or almost all, activities most of the day. As one of the two core cardinal symptoms of unipolar depression, anhedonia reflects impairment in positive valence systems, reward anticipation, and motivational consummation.
  • Depressed Mood (Item 2): Evaluates pervasive affective dysphoria, sadness, emotional emptiness, tearfulness, and feelings of profound hopelessness. It constitutes the second cardinal diagnostic feature required for syndromal diagnosis.
  • Sleep Disturbance (Item 3): Assesses vegetative dysregulation manifesting as insomnia (initial, middle, or terminal) or hypersomnia, reflecting circadian and neurochemical disruption.
  • Energy Deficit and Fatigue (Item 4): Measures subjective anergia, physical exhaustion, and diminished vitality unmitigated by rest, presenting frequently in primary care as physical complaints.
  • Appetite and Weight Alterations (Item 5): Captures bidirectional neurovegetative dysregulation manifested by significant appetite reduction (hyporexia) or compulsive overeating (hyperphagia).
  • Negative Self-Evaluation (Item 6): Assesses cognitive distortions centering on excessive guilt, feelings of failure, perceived inadequacy, and subjective sense of having let down oneself or family members, mapping directly onto Beck’s cognitive triad.
  • Executive and Attentional Dysfunctions (Item 7): Evaluates difficulties in sustained attention, concentration, working memory, and decision-making capacity during daily activities (e.g., reading print or following broadcast media).
  • Psychomotor Alterations (Item 8): Measures observable motor slowing (psychomotor retardation) or motor restlessness (psychomotor agitation), which often correlate with severe neurobiological depressive endophenotypes.
  • Suicidal Ideation and Self-Harm Thoughts (Item 9): Evaluates passive death wishes, active suicidal ideation, and intentional self-harm impulses, serving as an indispensable marker of clinical crisis and acute safety risk.

6. Theoretical Framework

The theoretical framework of the PHQ-9 integrates operational nosology, cognitive theories of depression, and measurement-based care (MBC) principles. Conceptually, it is grounded in the criterion-based diagnostic framework inaugurated by the St. Louis (Feighner) criteria, the Research Diagnostic Criteria (RDC) formulated by Spitzer, Endicott, and Robins, and ultimately the operationalized classification system codified in the Diagnostic and Statistical Manual of Mental Disorders.

This operational paradigm posits that mental disorders can be reliably identified by specifying concrete, observable behavioral and cognitive criteria evaluated across explicitly defined time windows and symptom thresholds. By formalizing these nine criteria into self-administered items, the PHQ-9 minimizes clinician-dependent diagnostic subjectivity and intra-rater variability. In parallel, the scale reflects Aaron T. Beck’s cognitive theory of depression, which demonstrates that depressogenic schemas, systematic cognitive distortions, and negative evaluations of the self, world, and future directly precipitate affective and neurovegetative manifestations.

From a psychometric measurement perspective, the PHQ-9 operationalizes the construct of depression under both Classical Test Theory (CTT) and Item Response Theory (IRT). Under CTT, the items are conceived as parallel reflections of a common underlying latent trait, where composite summation reflects overall severity. IRT modeling (specifically graded response models) has shown that items possess differential discrimination parameters ($lpha$) and difficulty thresholds ($eta$), with cognitive-affective items (e.g., anhedonia, depressed mood, feelings of failure) demonstrating the highest diagnostic information across moderate-to-severe regions of the latent depression trait ($ heta$).

7. Validity

The construct, criterion, convergent, and discriminant validities of the PHQ-9 have been substantiated across hundreds of validation studies in diverse clinical cohorts:

  • Criterion and Diagnostic Validity: In the landmark validation study by Kroenke, Spitzer, and Williams (2001) involving 6,000 patients across primary care and obstetrics/gynecology clinics, a PHQ-9 score $ge 10$ yielded a sensitivity of 88% and a specificity of 88% for major depression when compared against blind, structured psychiatric diagnostic interviews conducted by mental health professionals. Subsequent meta-analyses (e.g., Moriarty et al., 2015; Levis et al., 2019) synthesizing data across dozens of clinical studies comprising tens of thousands of participants established pooled sensitivity estimates between 0.80 and 0.85 and pooled specificity estimates between 0.85 and 0.90 for the $ge 10$ cutoff.
  • Convergent Validity: The PHQ-9 exhibits robust positive correlations with other validated depression instruments. It correlates strongly with the Beck Depression Inventory (BDI-II) ($r = 0.73$ to $0.84$), the Hamilton Depression Rating Scale (HAM-D) ($r = 0.71$ to $0.80$), and the General Health Questionnaire (GHQ-12) ($r = 0.65$ to $0.78$).
  • Construct and Discriminant Validity: PHQ-9 total scores correlate strongly with functional impairment and health-related quality of life metrics, showing strong negative correlations with the SF-20 / SF-36 Physical Functioning and Mental Health subscales ($r = -0.55$ to $-0.70$). Higher PHQ-9 scores reliably predict increased medical utilization, emergency visits, ambulatory encounters, and work absenteeism. Discriminant validity is supported by its ability to differentiate clinical major depression from non-depressive somatic complaints and generalized medical illnesses.
  • Predictive and Longitudinal Responsiveness: Longitudinal investigations establish that reductions in PHQ-9 scores correlate with clinical recovery. A decrease of $ge 5$ points is recognized as the minimal clinically important difference (MCID), while a post-treatment score of $<5$ indicates sustained clinical remission.

8. Reliability

The reliability profile of the PHQ-9 has been confirmed across clinical populations, age groups, and translated adaptations:

  • Internal Consistency: In the primary validation cohort ($N = 6,000$), the internal consistency of the PHQ-9 demonstrated a Cronbach’s alpha coefficient of $\alpha = 0.89$ in primary care samples and $\alpha = 0.86$ in obstetrics-gynecology settings. Subsequent large-scale investigations worldwide have reported alpha values consistently ranging between 0.84 and 0.91, indicating excellent item interrelatedness without item redundancy. McDonald’s omega ($\omega$) values typically match or exceed 0.88.
  • Test-Retest Reliability: Test-retest stability was evaluated by administering the questionnaire to patients at clinic check-in and re-administering it within 48 to 72 hours. The resulting intraclass correlation coefficient (ICC) and Pearson correlation coefficient were $r = 0.84$, demonstrating high temporal stability across brief reassessment intervals during which true depressive symptomatology remains largely unchanged.
  • Inter-Format Equivalence: Psychometric comparisons between self-administered paper-and-pencil forms, computerized/digital administration platforms, interactive voice response (IVR) telephone assessments, and clinician-read administrations have yielded cross-modal equivalence coefficients exceeding $r = 0.90$, demonstrating measurement stability regardless of administration format.

9. Factor Analysis

The structural dimensionality of the PHQ-9 has been extensively analyzed using exploratory factor analysis (EFA) and confirmatory factor analysis (CFA):

  • Unidimensional Model: The original validation framework supported an essentially unidimensional factor structure, wherein all nine items load significantly onto a single latent general depression dimension ($ heta$). Factor loadings for all nine items typically exceed 0.50, ranging from 0.52 (Item 8: psychomotor changes) to 0.82 (Item 2: depressed mood). Single-factor CFA models in normative primary care cohorts regularly yield acceptable fit indices: Comparative Fit Index (CFI)$ge 0.95$, Tucker-Lewis Index (TLI)$ge 0.94$, Root Mean Square Error of Approximation (RMSEA)$le 0.06$, and Standardized Root Mean Square Residual (SRMR)$le 0.04$.
  • Two-Factor (Somatic vs. Cognitive/Affective) Model: Several CFA investigations (e.g., Krause et al., 2010; Elhai et al., 2012) have identified improved model fit using a correlated two-factor model or a bifactor structure:
    • Cognitive/Affective Factor: Loaded primarily by Item 1 (anhedonia), Item 2 (depressed mood), Item 6 (worthlessness/guilt), Item 7 (concentration difficulties), and Item 9 (suicidal ideation).
    • Somatic/Vegetative Factor: Loaded primarily by Item 3 (sleep disturbance), Item 4 (fatigue/energy deficit), Item 5 (appetite changes), and Item 8 (psychomotor alterations).
  • Bifactor and General Factor Dominance: Bifactor modeling indicates that while the somatic and cognitive/affective group factors explain unique variance, the general depression factor accounts for the vast majority of common variance (Explained Common Variance [ECV] $> 0.75$, Omega Hierarchical [$\omega_h$] $> 0.80$). This justifies the conventional scoring and clinical interpretation of the PHQ-9 as an integrated composite score.
  • Measurement Invariance: Extensive multigroup CFA has documented full scalar and metric measurement invariance across gender, age groups, racial/ethnic backgrounds, and medical comorbidities, confirming that observed score differences reflect genuine differences in latent depression rather than item measurement bias.

10. Instrument / Measurement Tool

  • Instrument Type: Self-administered psychological assessment instrument / patient-reported outcome measure (PROM).
  • Administration Format: Paper-and-pencil questionnaire, digital web-based application, mobile tablet interface, or clinician-facilitated interview.
  • Target Population: Adults and adolescents aged 12 years and older (specifically with validated youth versions such as PHQ-A/PHQ-9 modified for adolescents).
  • Completion Time: Approximately 2 to 5 minutes.
  • Item Count: 9 diagnostic items (plus 1 optional supplementary global functional impairment item).
  • Response Scale: 4-point frequency scale: 0 = Not at all, 1 = Several days, 2 = More than half the days, 3 = Nearly every day.
  • Recall Window: Over the last 2 weeks.
  • Scoring Procedure: The global severity score is computed by summing the numerical ratings across all 9 individual items, resulting in a continuous total score ranging from 0 to 27.
  • Clinical Severity Thresholds:
    • 0–4: Minimal or no depression (monitoring; no treatment typically indicated).
    • 5–9: Mild depression (watchful waiting; psychoeducation; re-evaluation at follow-up).
    • 10–14: Moderate depression (recommended threshold for clinical intervention; psychotherapy or pharmacotherapy).
    • 15–19: Moderately severe depression (active clinical intervention warranted; combined pharmacotherapy and psychotherapy).
    • 20–27: Severe depression (immediate multimodal therapeutic intervention, specialty referral, and risk management).
  • Categorical Diagnostic Algorithm: A Major Depressive Episode is suggested if at least 5 of the 9 items are rated at $ge 2$ (“More than half the days”), with Item 9 counting positively if rated $ge 1$, and at least one of the positive symptoms is Item 1 (anhedonia) or Item 2 (depressed mood).
  • Critical Safety Item: Any non-zero score on Item 9 (“Thoughts that you would be better off dead or of hurting yourself in some way”) mandates immediate comprehensive clinical evaluation for suicide risk.

11. Permissions & Fee and Test Year

The Patient Health Questionnaire – 9 was officially published in its full psychometrically validated format in 2001 (Kroenke, Spitzer, & Williams, 2001), following preliminary clinical validation of the broader PRIME-MD PHQ in 1999 (Spitzer et al., 1999). It was developed by Drs. Robert L. Spitzer, Janet B. W. Williams, Kurt Kroenke, and colleagues at Columbia University and collaborating medical institutions.

Licensing and Royalty Status: The PHQ-9 is in the public domain. No copyright permissions, royalties, licensing fees, or formal authorizations are required to use, reproduce, modify, translate, or integrate the instrument into academic research, clinical medical records, electronic health record (EHR) systems, or commercial digital health applications. While Pfizer Inc. held the initial intellectual property copyright derived from grant sponsorship, the developers and Pfizer explicitly released the PHQ suite into the public commons to promote open, unhindered depression screening globally.

12. References

  • Beck, A. T., Steer, R. A., & Brown, G. K. (1996). Manual for the Beck Depression Inventory-II. Psychological Corporation. https://doi.org/10.1037/t00742-000
  • Elhai, J. D., Contractor, A. A., Tamburrino, M., Fine, T. H., Prescott, M. R., Shirley, E., Galea, S., & Calabrese, J. R. (2012). The factor structure of major depression symptoms: A test of four competing models using the Patient Health Questionnaire-9. Psychiatry Research, 197(1-2), 109–113. https://doi.org/10.1016/j.psychres.2011.12.046
  • Krause, J. S., Reed, K. S., & McArdle, J. J. (2010). Factor structure and predictive validity of the Patient Health Questionnaire-9 in persons with spinal cord injury. Archives of Physical Medicine and Rehabilitation, 91(7), 1018–1025. https://doi.org/10.1016/j.apmr.2010.03.018
  • Kroenke, K., & Spitzer, R. L. (2002). The PHQ-9: A new depression diagnostic and severity measure. Psychiatric Annals, 32(9), 509–515. https://doi.org/10.3928/0048-5713-20020901-06
  • Kroenke, K., Spitzer, R. L., & Williams, J. B. W. (2001). The PHQ-9: Validity of a brief depression severity measure. Journal of General Internal Medicine, 16(9), 606–613. https://doi.org/10.1046/j.1525-1497.2001.016009606.x
  • Levis, B., Benedetti, A., & Thombs, B. D. (2019). Accuracy of Patient Health Questionnaire-9 (PHQ-9) for screening to detect major depression: An individual participant data meta-analysis. BMJ, 365, l1476. https://doi.org/10.1136/bmj.l1476
  • Moriarty, A. S., Gilbody, S., McMillan, D., & Manea, L. (2015). Screening and case finding for major depressive disorder using the Patient Health Questionnaire-9: A meta-analysis. General Hospital Psychiatry, 37(6), 567–576. https://doi.org/10.1016/j.genhosppsych.2015.06.012
  • Spitzer, R. L., Kroenke, K., & Williams, J. B. W. (1999). Validation and utility of a self-report version of PRIME-MD: The PHQ primary care study. JAMA, 282(18), 1737–1744. https://doi.org/10.1001/jama.282.18.1737

13. Items of the Scale

Below are the authentic scale items in their original language as published in the standard psychometric validation studies, without modification or translation to preserve instrument validity and reliability:

Over the last 2 weeks, how often have you been bothered by any of the following problems?

Response Scale: 4-point frequency scale: 0 = Not at all, 1 = Several days, 2 = More than half the days, 3 = Nearly every day

  1. Little interest or pleasure in doing things
  2. Feeling down, depressed, or hopeless
  3. Trouble falling or staying asleep, or sleeping too much
  4. Feeling tired or having little energy
  5. Poor appetite or overeating
  6. Feeling bad about yourself — or that you are a failure or have let yourself or your family down
  7. Trouble concentrating on things, such as reading the newspaper or watching television
  8. Moving or speaking so slowly that other people could have noticed? Or the opposite — being so fidgety or restless that you have been moving around a lot more than usual
  9. Thoughts that you would be better off dead or of hurting yourself in some way

Rate This Scale

5.0 / 5 1 vote

Cite This Article

memjavad (2026, September 4). Patient Health Questionnaire – 9 (PHQ-9). PSYCHOLOGICAL DATABASE. https://en.arabpsychology.com/scales/patient-health-questionnaire-9-phq-9/
memjavad. “Patient Health Questionnaire – 9 (PHQ-9).” PSYCHOLOGICAL DATABASE, 4 September 2026, https://en.arabpsychology.com/scales/patient-health-questionnaire-9-phq-9/.
memjavad. “Patient Health Questionnaire – 9 (PHQ-9).” PSYCHOLOGICAL DATABASE. September 4, 2026. https://en.arabpsychology.com/scales/patient-health-questionnaire-9-phq-9/.