Clinical AssessmentHealth PsychologyPsychometrics

36-Item Short Form Health Survey

A definitive, comprehensive psychometric guide to the 36-Item Short Form Health Survey (SF-36 / RAND-36), detailing its theoretical framework, structural validity, subscale architecture, reliability, and full authentic administration items.

memjavad
PUBLISHED
Scientifically Reviewed · Dr. Marwa Abd-Alazim · September 12, 2026
Medically & Scientifically Reviewed Verified: September 12, 2026
Dr. Marwa Abd-Alazim Ph.D.
Professor of Psychology University of Kerbala
Review Criteria & Clinical Standards

This content undergoes rigorous scientific peer-review and medical editorial standards at Arab Psychology Network to ensure clinical accuracy, validity, and compliance with evidence-based guidelines from leading psychological and healthcare authorities (APA / WHO).

Abstract

The 36-Item Short Form Health Survey (SF-36), originally developed within the monumental Medical Outcomes Study (MOS) by John E. Ware, Jr. and Cathy Donald Sherbourne at the RAND Corporation, stands as the most widely implemented, empirically scrutinized, and validated generic patient-reported outcome measure (PROM) in international clinical research and public health surveillance. Comprising 36 items designed for self-administration, computer-assisted evaluation, or clinical interview, the instrument captures eight distinct dimensions of health-related quality of life (HRQoL): Physical Functioning (10 items), Role Limitations due to Physical Health (4 items), Bodily Pain (2 items), General Health Perceptions (5 items), Vitality/Energy/Fatigue (4 items), Social Functioning (2 items), Role Limitations due to Emotional Problems (3 items), and Emotional Well-being/Mental Health (5 items), alongside a single unscaled item documenting perceived health transitions over a one-year interval. Item responses utilize varied polytomous and dichotomous response scales (ranging from 2-point dichotomous formats to 3-point, 5-point, and 6-point graded categories), which are systematically coded and linearly transformed into continuous norm-based scales ranging from 0 (worst possible health state) to 100 (optimal health state). Furthermore, through orthogonal and oblique factor-analytic derivation, two standardized higher-order summary scales—the Physical Component Summary (PCS) and the Mental Component Summary (MCS)—are derived to facilitate parsimonious group comparisons and clinical trial outcome modeling. Extensive psychometric evaluations across hundreds of diverse clinical cohorts and general population samples confirm remarkable internal consistency (Cronbach’s alpha coefficients consistently exceeding .80 across subscales, and frequently exceeding .90 for physical functioning), high test-retest reliability (intraclass correlation coefficients of .75 to .90 across 2- to 4-week intervals), and robust construct, convergent, discriminant, and predictive validity across chronic physical conditions, psychiatric morbidities, and health-economic longitudinal trajectories.

Keywords

SF-36, RAND-36, Health-Related Quality of Life, Patient-Reported Outcome Measures, Psychometrics, Physical Component Summary, Mental Component Summary, Medical Outcomes Study, Construct Validity, Factor Analysis, Health Status Measurement

Authors

The foundational conceptualization and item generation of the SF-36 occurred under the auspices of the Health Program at the RAND Corporation (Santa Monica, California) in conjunction with the multi-center Medical Outcomes Study during the late 1980s. The primary investigators credited with the definitive operationalization, psychometric testing, and empirical dissemination of the instrument are:

  • John E. Ware, Jr., Ph.D.: Psychometrician, Senior Scientist at the Institute for the Improvement of Medical Care and Health, New England Medical Center, and Professor of Public Health at the University of Massachusetts Medical School and Harvard T.H. Chan School of Public Health. Dr. Ware led the conceptual structuring of health status assessment in the MOS and spearheaded subsequent versions, including the commercialized SF-36v2 via QualityMetric Incorporated.
  • Cathy Donald Sherbourne, Ph.D.: Senior Behavioral Scientist at the RAND Corporation, whose foundational empirical investigations into social functioning, vitality, and psychiatric distress in primary care cohorts directly framed the instrument’s multi-attribute dimensional taxonomy.
  • Anita L. Stewart, Ph.D.: Senior Researcher at the Institute for Health & Aging, University of California, San Francisco (UCSF), and RAND Corporation, central to the operational design of the original MOS battery and physical functioning subscales.
  • Neil K. Aaronson, Ph.D.: Division of Psychosocial Research and Epidemiology, The Netherlands Cancer Institute (NKI-AVL), Amsterdam. Dr. Aaronson served as the lead coordinator for the Dutch cultural adaptation, validation, and standard population norming of the SF-36/RAND-36, under the auspices of the International Quality of Life Assessment (IQOLA) Project.

Purpose

The primary purpose of the 36-Item Short Form Health Survey is to provide a standardized, psychometrically robust, and comprehensive measure of health-related quality of life that can be universally deployed across adult populations regardless of underlying diagnosis, sociodemographic variation, or treatment setting. Prior to the development of the SF-36, epidemiological and health services research relied heavily on either disease-specific instruments (which precluded direct comparative effectiveness evaluations across disparate clinical conditions) or cumbersome multidimensional batteries comprising hundreds of questions, such as the Sickness Impact Profile (SIP) or the Nottingham Health Profile (NHP). The SF-36 was systematically engineered to bridge this methodological divide by optimizing clinical brevity—averaging five to ten minutes for complete administration—without sacrificing the psychometric rigor, breadth, and discriminatory capacity inherent in comprehensive health status batteries.

In clinical trials, the SF-36 serves as a sensitive secondary or primary clinical endpoint, enabling investigators to quantify subtle therapeutic benefits, therapeutic adverse effects, and longitudinal trajectories in functional status following pharmacological, surgical, or behavioral interventions. In routine clinical practice, the tool facilitates systematic patient-reported monitoring, allowing clinicians to screen for silent functional impairments, monitor disease progression in chronic conditions such as rheumatoid arthritis, congestive heart failure, chronic obstructive pulmonary disease (COPD), and major depressive disorder, and adjust interventions based on subjective burden. In population health surveillance and health economics, the SF-36 and its derivation algorithms (such as the SF-6D utility index) are utilized to establish general population norms, compute Quality-Adjusted Life Years (QALYs), benchmark hospital and health-system performance, and inform health technology assessment (HTA) policy decisions globally.

Psychological Construct

The SF-36 measures a multifaceted, multidimensional construct of Health-Related Quality of Life (HRQoL), operationalized as the extent to which an individual’s physical, psychological, and social functioning and well-being are influenced by subjective perceptions of health, illness, or treatment. Drawing directly from the classic definition of health promulgated by the World Health Organization—namely, a complete state of physical, mental, and social well-being and not merely the absence of disease or infirmity—the SF-36 maps this construct into eight discrete, lower-order dimensions and two higher-order meta-domains:

1. Physical Functioning (PF; 10 items)

This subscale assesses the presence and severity of physical limitations across a continuous gradient of mobility and daily personal activities. The items range from severe limitations affecting basic survival self-care (e.g., bathing or dressing oneself) to moderate demands (e.g., moving a table, pushing a vacuum cleaner, climbing a single flight of stairs, walking several blocks) up to strenuous physical exertion (e.g., running, lifting heavy objects, participating in vigorous sports). Low scores reflect marked functional dependency and profound physical restrictions, whereas high scores indicate the unrestricted performance of all vigorous physical demands without health-induced inhibition.

2. Role Limitations due to Physical Health (RP; 4 items)

This dimension measures the degree to which an individual’s physical health problems interfere with their occupational functioning, domestic responsibilities, or daily routines. It quantifies four specific behavioral manifestations: having to cut down on time spent on work or tasks, achieving less than desired, experiencing limitations in the types of tasks performed, and encountering heightened difficulty in performing customary work. Low scores reflect severe career or domestic paralysis induced by physical illness.

3. Bodily Pain (BP; 2 items)

This subscale captures both the sensory intensity of bodily pain experienced over the preceding four weeks (ranging from none to very severe) and the secondary functional interference such pain causes with normal occupational and domestic duties. It operates as a distinct somatic indicator that directly cross-cuts physical performance and psychological distress.

4. General Health Perceptions (GH; 5 items)

This dimension integrates subjective appraisals of current health status, personal susceptibility to illness (e.g., “I seem to get sick a little easier than other people”), resistance against disease deterioration (“I expect my health to get worse”), and overall health excellence relative to peers. Unlike purely objective physical performance scales, this construct evaluates the cognitive-affective meaning and internal benchmark an individual attributes to their biological state.

5. Vitality / Energy / Fatigue (VT; 4 items)

Capturing a continuous affective-somatic spectrum, this subscale evaluates feelings of energy, pep, exhaustion, and fatigue experienced over the previous four weeks. As a bipolar construct, it discriminates between debilitating, pervasive somatic weariness and vibrant energetic engagement, functioning psychometrically as a bridge linking pure physical capability with internal psychological drive.

6. Social Functioning (SF; 2 items)

This subscale quantifies the degree and frequency to which physical health impairments or emotional difficulties disrupt normal interpersonal interactions, community activities, and social networks with family, friends, neighbors, and peer groups. It isolates social isolation and relationship strain secondary to poor health.

7. Role Limitations due to Emotional Problems (RE; 3 items)

Analogous to the RP subscale, this construct evaluates occupational and daily role deficits that stem specifically from psychological distress, affective disturbance, or anxiety. It isolates behaviors such as reducing time spent on tasks, accomplishing less than desired, and executing duties with compromised attention, precision, or quality as a consequence of internal emotional conflict.

8. Emotional Well-Being / Mental Health (MH; 5 items)

This subscale comprehensively samples four prominent psychological dimensions: generalized anxiety (nervousness), depression (feeling downhearted, blue, and down in the dumps), positive affect (happiness, peacefulness), and behavioral-emotional control over the past four weeks. Elevated scores signify sustained psychological resilience, positive hedonic tone, and calm, whereas depressed scores denote severe psychological distress, dysphoria, and generalized affective demoralization.

Higher-Order Summaries: PCS and MCS

Through empirical factor analysis, the eight scales coalesce into two overarching constructs: the Physical Component Summary (PCS), heavily weighted by Physical Functioning, Role-Physical, and Bodily Pain; and the Mental Component Summary (MCS), predominantly driven by Emotional Well-being, Role-Emotional, and Social Functioning, with Vitality and General Health displaying substantial intermediate cross-loadings across both higher-order vectors.

Theoretical Framework

The theoretical architecture underpinning the SF-36 rests upon the intersection of George Engel’s Biopsychosocial Model and classical psychometric Classical Test Theory (CTT), augmented by contemporary Item Response Theory (IRT) principles. In 1977, Engel postulated that understanding clinical illness requires moving beyond the reductionist, purely biomedical paradigm to examine the dynamic interplay between biological pathophysiology, psychological vulnerability, and socioeconomic/environmental contexts. The SF-36 operationalizes this paradigm by conceptualizing health not as an isolated laboratory biomarker, but as an integrated continuum of biological capacity, functional task performance, psychological affect, and social role execution.

During the Medical Outcomes Study, Ware and colleagues developed a hierarchical conceptual model of health. At the most fundamental biological tier lies physiological pathology (e.g., cellular impairment, inflammatory cascade). This pathology translates into personal awareness of symptoms (e.g., pain, dyspnea, nausea, fatigue). Symptom awareness induces functional limitations in physical, cognitive, or interpersonal capacity. When sustained, functional limitations constrain role performance across family, work, and community settings, ultimately shaping the individual’s macro-level cognitive evaluation of their subjective well-being and general health expectations. The SF-36 intentionally focuses measurement at the levels of symptoms, functional limitations, role enactments, and global cognitive appraisals, bypassing disease-specific mechanisms to create an interchangeable metric across medical specialties.

Psychometrically, the construct assumes that latent dimensions of physical and mental health can be reliably isolated using a multi-trait scaling methodology. Ware utilized cumulative scaling models and Likert’s technique of summated ratings, asserting that individual items within a subscale are linear manifestations of a single underlying latent trait. Later IRT calibrations (using Graded Response Models) confirmed that items across the physical functioning continuum possess monotonic difficulty progressions (e.g., bathing oneself exhibits a much higher threshold of impairment than running a mile), providing empirical justification for the tool’s wide measurement span across both healthy and severely compromised clinical populations.

Validity

The validity of the SF-36 has been investigated across hundreds of diverse clinical cohorts, healthy demographic cross-sections, and cultural translations worldwide.

Content and Face Validity

Content validity was established during the MOS through comprehensive reviews of existing health status inventories, rigorous cognitive debriefing interviews with chronically ill patients, and consensus panels composed of epidemiologists, clinicians, and psychometricians. The final 36 items represent an optimal operationalization of the core domains identified by the WHO, showing universal face validity across diverse demographic backgrounds.

Construct, Convergent, and Discriminant Validity

Construct validity was formally substantiated using the Multitrait-Multimethod Matrix approach during the initial validation studies by Ware and Sherbourne (1992) and the International Quality of Life Assessment (IQOLA) project (Aaronson et al., 1998; Ware et al., 1995). Item-scale convergent validity is exceptionally high: item-total correlations within their designated subscales routinely exceed the standard .40 psychometric threshold, typically falling between .55 and .82 (corrected for item-overlap). Discriminant validity is affirmed by the fact that correlations between individual items and their parent subscales are significantly higher (often by two or more standard deviations) than their secondary correlations with outside subscales.

Convergent validity has been repeatedly documented against legacy clinical instruments. The Physical Functioning subscale correlates strongly ($r > .70$) with the Health Assessment Questionnaire (HAQ) Disability Index, the Karnofsky Performance Scale, and measured physical performance tests (e.g., the 6-Minute Walk Test). Conversely, the Mental Health/Emotional Well-being subscale demonstrates strong correlations ($r = -.75$ to $-.85$) with established psychiatric depression and anxiety inventories, including the Beck Depression Inventory (BDI) and the Hospital Anxiety and Depression Scale (HADS), while maintaining low correlations ($r < .30$) with pure physical parameters such as forced expiratory volume ($FEV_1$) or left ventricular ejection fraction.

Criterion and Known-Groups Validity

Known-groups validity has been extensively demonstrated across medical disciplines. The SF-36 discriminates accurately between populations with differing clinical severity. For example, individuals with heart failure, severe osteoarthritis, or advanced COPD exhibit significantly depressed Physical Functioning and PCS scores (often 1.5 to 2.5 standard deviations below population means), whereas their Mental Component Summary scores frequently parallel or only slightly trail normative benchmarks. Conversely, individuals diagnosed with major depressive episodes or generalized anxiety disorder demonstrate profound deficits on the Mental Health, Role-Emotional, and MCS scores, while displaying preserved Physical Functioning scores. Longitudinal predictive validity has also been verified: low baseline PCS scores prospectively predict three-to-five-year mortality, hospitalization frequency, and permanent disability claims in elderly cohorts, independent of objective comorbid biological pathology.

Reliability

The psychometric reliability of the SF-36 has been confirmed through repeated evaluations across cross-sectional surveys, longitudinal population registries, and multi-site randomized controlled clinical trials.

Internal Consistency Reliability

Across validation studies spanning thousands of general population respondents and clinical cohorts, internal consistency coefficients (Cronbach’s alpha) uniformly exceed the acceptable scientific benchmark of .70 for group comparisons, and consistently surpass the stringent .90 threshold required for individual-level clinical decision-making. Standard empirical parameters observed across published literature (Ware et al., 1993; Aaronson et al., 1998; Gandek et al., 1998) include:

  • Physical Functioning (PF): $\alpha = .90 – .94$
  • Role Limitations – Physical (RP): $\alpha = .84 – .91$
  • Bodily Pain (BP): $\alpha = .78 – .88$
  • General Health (GH): $\alpha = .78 – .85$
  • Vitality (VT): $\alpha = .84 – .88$
  • Social Functioning (SF): $\alpha = .76 – .85$
  • Role Limitations – Emotional (RE): $\alpha = .80 – .88$
  • Emotional Well-Being / Mental Health (MH): $\alpha = .84 – .90$

The composite higher-order summary scores, PCS and MCS, exhibit exceptional internal consistency, with reliability coefficients calculated via Mosier’s composite reliability formula typically hovering between .92 and .95.

Test-Retest Reliability and Measurement Error

Test-retest stability has been demonstrated across stable clinical and healthy cohorts re-evaluated at two-week to one-month intervals. Intraclass Correlation Coefficients (ICC) range from .70 for scales with fewer items (e.g., Social Functioning and Bodily Pain) to .85 to .92 for Physical Functioning, General Health, and the composite PCS/MCS indices. The Standard Error of Measurement (SEM) is modest, permitting the identification of the Minimal Clinically Important Difference (MCID). Across diverse chronic conditions, the MCID has been empirically determined to be approximately 3 to 5 points on the 0–100 transformed subscale metrics, and 2 to 3 points on the standardized norm-based PCS and MCS summary scales (where the general population mean is pegged to 50 with an SD of 10).

Factor Analysis

The internal structural validity of the SF-36 has been extensively mapped using both Exploratory Factor Analysis (EFA) and Confirmatory Factor Analysis (CFA) across diverse linguistic, geographic, and disease cohorts.

Exploratory Factor Analysis (EFA)

In original factor analyses conducted by Ware and colleagues (1994) using principal components analysis with varimax (orthogonal) and promax (oblique) rotations, a clear two-factor structure emerged from the correlation matrix of the eight primary subscales. These two factors accounted for approximately 60% to 65% of the total variance across general and clinical populations:

  • Factor 1: Physical Health Vector. Dominated by high factor loadings from Physical Functioning (.85 to .88), Role-Physical (.78 to .82), and Bodily Pain (.74 to .79).
  • Factor 2: Mental Health Vector. Dominated by substantial factor loadings from Emotional Well-Being / Mental Health (.84 to .89), Role-Emotional (.77 to .82), and Social Functioning (.70 to .75).
  • Dual-Loading Subscales. Vitality, General Health, and Social Functioning consistently exhibit moderate cross-loadings (.35 to .55) on both physical and mental factors, providing structural validation for their role as systemic cross-domain bridges.

Confirmatory Factor Analysis (CFA)

Confirmatory factor analytic investigations have tested hierarchical and correlated-factor structural models. A correlated two-factor higher-order model (allowing physical and mental constructs to covary, typically between $r = .40$ and $.65$) generally achieves superior goodness-of-fit indices compared to strictly orthogonal models:

  • Comparative Fit Index (CFI): $ge .94 – .97$
  • Tucker-Lewis Index (TLI): $ge .93 – .96$
  • Root Mean Square Error of Approximation (RMSEA): $le .045 – .062$
  • Standardized Root Mean Square Residual (SRMR): $le .040 – .055$

Cross-national multigroup CFA conducted by the IQOLA project across 15 nations established structural invariance (metric and scalar invariance), confirming that the underlying factor loading patterns and intercept structures remain uniform across distinct languages and cultural contexts.

Instrument / Measurement Tool

  • Instrument Name: 36-Item Short Form Health Survey (SF-36 / RAND-36).
  • Construct Measured: Health-Related Quality of Life (HRQoL); multidimensional generic physical and mental health status.
  • Test Type: Patient-Reported Outcome Measure (PROM); self-report questionnaire, structured clinician interview, or digital computer-adaptive test.
  • Target Population: Adults (aged 18 and older) and older adults across healthy community populations and diverse medical cohorts. (Youth versions, such as the SF-10, exist separately).
  • Administration Time: Approximately 5 to 10 minutes.
  • Number of Items: 36 items (35 items mapping into eight scaled health domains, plus 1 self-standing unscaled health transition item).
  • Subscale Breakdown:
    • Physical Functioning (PF): Items 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 (10 items)
    • Role Limitations due to Physical Health (RP): Items 13, 14, 15, 16 (4 items)
    • Bodily Pain (BP): Items 21, 22 (2 items)
    • General Health Perceptions (GH): Items 1, 33, 34, 35, 36 (5 items)
    • Vitality / Energy / Fatigue (VT): Items 23, 27, 29, 31 (4 items)
    • Social Functioning (SF): Items 20, 32 (2 items)
    • Role Limitations due to Emotional Problems (RE): Items 17, 18, 19 (3 items)
    • Emotional Well-Being / Mental Health (MH): Items 24, 25, 26, 28, 30 (5 items)
    • Reported Health Transition: Item 2 (1 item, evaluated independently)
  • Authentic Response Scale: Varies by item section: 5-point rating (1=Excellent to 5=Poor; 1=Much better now than one year ago to 5=Much worse now than one year ago; 1=Definitely true to 5=Definitely false); 3-point rating (1=Yes, limited a lot, 2=Yes, limited a little, 3=No, not limited at all); 2-point dichotomous (1=Yes, 2=No); 6-point rating (1=None to 6=Very severe; 1=Not at all to 5=Extremely; 1=All of the time to 6=None of the time)
  • Scoring and Transformation Procedures:
    • Step 1: Recoding of Items. Items 1, 2, 20, 22, 34, and 36, along with negatively framed items in vitality and emotional well-being (items 24, 25, 28, 29, 31), undergo reverse scoring or algebraic recalibration so that higher raw numbers uniformly reflect more favorable health states.
    • Step 2: Computation of Raw Scale Scores. Summing calibrated scores across items within each discrete subscale. Missing values are imputed if the respondent answered at least 50% of the items within that specific subscale (person-mean imputation).
    • Step 3: Linear Transformation (0 to 100 Scale). Raw scores are converted to a standardized 0 to 100 scale using the linear formula:
      $$\text{Transformed Score} = \left( \frac{\text{Actual Raw Score} – \text{Lowest Possible Raw Score}}{\text{Possible Raw Score Range}} \right) \times 100$$
      where 0 indicates the worst possible subjective health state and 100 indicates the optimal, completely unimpaired health state.
    • Step 4: Derivation of Summary Measures (PCS and MCS). Transformed scores for the eight subscales are standardized into Z-scores using reference general population means and standard deviations, multiplied by orthogonal or oblique factor score coefficients, summed, and transformed into norm-based T-scores ($Mean = 50, SD = 10$).
    • RAND-36 vs. SF-36 Distinction: The RAND-36 uses public domain scoring algorithms where items 21 and 22 (Bodily Pain) and item 20 (Social Functioning) are scored slightly differently than in commercial QualityMetric SF-36 algorithms; however, the correlation between scores derived from the two methods exceeds .99 across all domains.

Permissions & Fee and Test Year

The 36-Item Short Form Health Survey was first finalized and published in 1992 following pilot deployment in 1990 under the Medical Outcomes Study. Because the initial development was supported by public research grants at the RAND Corporation, the RAND 36-Item Health Survey 1.0 was placed directly in the public domain. Researchers and clinicians may utilize the RAND-36 item set and its corresponding scoring rules free of licensing fees, royalties, or copyright restrictions for non-commercial academic research, public health surveillance, and clinical care, provided appropriate bibliographic citation is accorded to Ware and Sherbourne (1992).

Conversely, the commercial designations “SF-36®” and “SF-36v2®” are registered trademarks owned by QualityMetric Incorporated (an Optum company). SF-36 Version 2 (SF-36v2), which introduced minor grammatical rewordings, improved formatting, and extended the 2-point dichotomous response scales of the RP and RE subscales into 5-point rating scales to resolve historical floor and ceiling effects, requires formal licensing, user agreements, and commercial per-administration or enterprise royalty fees managed through QualityMetric. For non-funded research and standard independent academic research where commercial licensing is prohibitive, the original public-domain RAND-36 instrument remains widely used and scientifically comparable.

References

  • Aaronson, N. K., Muller, M., Cohen, P. D., Essink-Bot, M. L., Fekkes, M., Sanderman, R., Sprangers, M. A., te Velde, A., & Verrips, E. (1998). Translation, validation, and norming of the Dutch language version of the SF-36 Health Survey in community and chronic disease populations. Journal of Clinical Epidemiology, 51(11), 1055–1068. https://doi.org/10.1016/s0895-4356(98)00097-3
  • Engel, G. L. (1977). The need for a new medical model: A challenge for biomedicine. Science, 196(4286), 129–136. https://doi.org/10.1126/science.847460
  • Gandek, B., Ware, J. E., Aaronson, N. K., Alonso, J., Apolone, G., Bjorner, J., Brazier, J., Bullinger, M., Kaasa, S., Leplege, A., & Sullivan, M. (1998). Tests of data quality, scaling assumptions, and reliability of the SF-36 Health Survey across nine countries: Results from the IQOLA Project. Journal of Clinical Epidemiology, 51(11), 1149–1158. https://doi.org/10.1016/s0895-4356(98)00106-1
  • McHorney, C. A., Ware, J. E., & Raczek, A. E. (1993). The MOS 36-Item Short-Form Health Survey (SF-36): II. Psychometric and clinical tests of validity in measuring physical and mental health constructs. Medical Care, 31(3), 247–263. https://doi.org/10.1097/00005650-199303000-00006
  • Stewart, A. L., Hays, R. D., & Ware, J. E. (1988). The MOS short-form general health survey: Reliability and validity in a patient population. Medical Care, 26(7), 724–735. https://doi.org/10.1097/00005650-198807000-00007
  • Tarlov, A. R., Ware, J. E., Greenfield, S., Nelson, E. C., Perrin, E., & Zubkoff, M. (1989). The Medical Outcomes Study: An approach to evaluating the results of medical practice. JAMA, 262(7), 925–930. https://doi.org/10.1001/jama.1989.03430070073033
  • Ware, J. E., & Gandek, B. (1998). Overview of the SF-36 Health Survey and the International Quality of Life Assessment (IQOLA) Project. Journal of Clinical Epidemiology, 51(11), 903–912. https://doi.org/10.1016/s0895-4356(98)00081-x
  • Ware, J. E., Kosinski, M., & Keller, S. D. (1994). SF-36 physical and mental health summary scales: A user’s manual. The Health Institute, New England Medical Center.
  • Ware, J. E., & Sherbourne, C. D. (1992). The MOS 36-item short-form health survey (SF-36): I. Conceptual framework and item selection. Medical Care, 30(6), 473–483. https://doi.org/10.1097/00005650-199206000-00002
  • Ware, J. E., Snow, K. K., Kosinski, M., & Gandek, B. (1993). SF-36 Health Survey: Manual and interpretation guide. The Health Institute, New England Medical Center.

13. Items of the Scale (Questionnaire)

Below are the authentic scale items in their original language as published in the standard psychometric validation studies, without modification or translation to preserve instrument validity and reliability:
Instructions / Directions: This survey asks for your views about your health. This information will help keep track of how you feel and how well you are able to do your usual activities. Answer every question by selecting the indicated answer. If you are unsure about how to answer, please give the best answer you can.
Response Scale: Varies by item section: 5-point rating (1=Excellent to 5=Poor; 1=Much better now than one year ago to 5=Much worse now than one year ago; 1=Definitely true to 5=Definitely false); 3-point rating (1=Yes, limited a lot, 2=Yes, limited a little, 3=No, not limited at all); 2-point dichotomous (1=Yes, 2=No); 6-point rating (1=None to 6=Very severe; 1=Not at all to 5=Extremely; 1=All of the time to 6=None of the time)
Scoring / Reverse Items: Items are recoded per the MOS/RAND SF-36 scoring manual so that higher scores represent better health (0 to 100 scale). Items 1, 2, 20, 22, 34, and 36 are recoded. Subscales include: Physical Functioning (items 3-12), Role Limitations due to Physical Health (items 13-16), Role Limitations due to Emotional Problems (items 17-19), Energy/Fatigue (items 23, 27, 29, 31), Emotional Well-being (items 24, 25, 26, 28, 30), Social Functioning (items 20, 32), Bodily Pain (items 21, 22), and General Health (items 1, 33, 34, 35, 36). Item 2 measures reported health transition.
1

In general, would you say your health is:
2

Compared to one year ago, how would you rate your health in general now?
3

Vigorous activities, such as running, lifting heavy objects, participating in strenuous sports
4

Moderate activities, such as moving a table, pushing a vacuum cleaner, bowling, or playing golf
5

Lifting or carrying groceries
6

Climbing several flights of stairs
7

Climbing one flight of stairs
8

Bending, kneeling, or stooping
9

Walking more than a mile
10

Walking several blocks
11

Walking one block
12

Bathing or dressing yourself
13

Cut down the amount of time you spent on work or other activities (due to your physical health)
14

Accomplished less than you would like (due to your physical health)
15

Were limited in the kind of work or other activities (due to your physical health)
16

Had difficulty performing the work or other activities (for example, it took extra effort) (due to your physical health)
17

Cut down the amount of time you spent on work or other activities (due to emotional problems)
18

Accomplished less than you would like (due to emotional problems)
19

Didn't do work or other activities as carefully as usual (due to emotional problems)
20

During the past 4 weeks, to what extent has your physical health or emotional problems interfered with your normal social activities with family, friends, neighbors, or groups?
21

How much bodily pain have you had during the past 4 weeks?
22

During the past 4 weeks, how much did pain interfere with your normal work (including both work outside the home and housework)?
23

Did you feel full of pep?
24

Have you been a very nervous person?
25

Have you felt so down in the dumps that nothing could cheer you up?
26

Have you felt calm and peaceful?
27

Did you have a lot of energy?
28

Have you felt downhearted and blue?
29

Did you feel worn out?
30

Have you been a happy person?
31

Did you feel tired?
32

During the past 4 weeks, how much of the time has your physical health or emotional problems interfered with your social activities (like visiting with friends, relatives, etc.)?
33

I seem to get sick a little easier than other people
34

I am as healthy as anybody I know
35

I expect my health to get worse
36

My health is excellent

Rate This Scale

5.0 / 5 1 vote

Cite This Article

memjavad (2026, September 12). 36-Item Short Form Health Survey. PSYCHOLOGICAL DATABASE. https://en.arabpsychology.com/scales/36-item-short-form-health-survey/
memjavad. “36-Item Short Form Health Survey.” PSYCHOLOGICAL DATABASE, 12 September 2026, https://en.arabpsychology.com/scales/36-item-short-form-health-survey/.
memjavad. “36-Item Short Form Health Survey.” PSYCHOLOGICAL DATABASE. September 12, 2026. https://en.arabpsychology.com/scales/36-item-short-form-health-survey/.