Clinical AssessmentPsychometricsRehabilitation Psychology

Goal Attainment Scale

Goal Attainment Scaling (GAS) is an individualized psychometric methodology formulated by Kiresuk and Sherman to establish and evaluate personalized treatment goals using a standardized 5-point ordinal scale transformed into a parametric composite T-score.

memjavad
PUBLISHED
Scientifically Reviewed · Dr. Marwa Abd-Alazim · September 12, 2026
Medically & Scientifically Reviewed Verified: September 12, 2026
Dr. Marwa Abd-Alazim Ph.D.
Professor of Psychology University of Kerbala
Review Criteria & Clinical Standards

This content undergoes rigorous scientific peer-review and medical editorial standards at Arab Psychology Network to ensure clinical accuracy, validity, and compliance with evidence-based guidelines from leading psychological and healthcare authorities (APA / WHO).

1. Abstract

Goal Attainment Scaling (GAS) is an individualized, criterion-referenced psychometric evaluation methodology originally conceived by Thomas J. Kiresuk and Robert E. Sherman in 1968. Developed to evaluate treatment outcomes within comprehensive community mental health services, the methodology has since achieved ubiquitous adoption across physical therapy, occupational therapy, neurorehabilitation, geriatric care, pediatric early intervention, and palliative medicine. Rather than administering a static set of normative inventory items, GAS operationalizes an idiographic paradigm wherein patient-specific functional objectives are established a priori through collaborative clinical negotiation. Each identified therapeutic goal is calibrated across a standard five-point ordinal scale of predicted attainment: -2 representing the baseline or significantly worse performance, -1 denoting less than expected progress, 0 representing the expected level of outcome following intervention, +1 indicating greater than expected outcome, and +2 denoting much greater than expected outcome.

By standardizing discrete, clinically diverse goals into a mathematically uniform metric, GAS facilitates parametric aggregation via a standardized composite T-score, characterized theoretically by a mean of 50 and a standard deviation of 10. This structural versatility allows diverse rehabilitation interventions—spanning pediatric neurodevelopment to adult neurodegenerative conditions such as Parkinson’s disease—to be systematically evaluated. Psychometric appraisals confirm that GAS displays exceptional responsiveness and clinical sensitivity to subtle, ecologically valid functional changes that standard standardized psychometric batteries frequently fail to detect. Concurrently, inter-rater reliability coefficients historically range between r = 0.65 and 0.93 when explicit, SMART-compliant operational definitions are established, and construct validity is confirmed via convergent correlations (ranging from r = 0.40 to 0.76) with established disability and functional independence inventories.

2. Keywords

Goal Attainment Scaling, GAS, Kiresuk and Sherman, individualized outcome measures, psychometrics, rehabilitation psychology, functional outcomes, SMART goals, clinical measurement, T-score transformation

3. Authors

Goal Attainment Scaling was originally formulated and published by:

  • Thomas J. Kiresuk, Ph.D. — Former Director of Research at the Hennepin County Mental Health Center and Chief Clinical Psychologist, Minneapolis, Minnesota; Professor Emeritus in the Department of Psychology and Health Care Psychology, University of Minnesota Medical School. Dr. Kiresuk was a primary pioneer in program evaluation methodologies and health services research.
  • Robert E. Sherman, Ph.D. — Former Chief Biostatistician, Hennepin County Mental Health Center, Minneapolis, Minnesota; Research Associate, Department of Psychiatry, University of Minnesota. Dr. Sherman developed the mathematical algorithm transforming individualized ordinal goal levels into standardized composite parametric T-scores.

Subsequent adaptations and clinical translation frameworks include:

  • Royal Dutch Society for Physical Therapy (Koninklijk Nederlands Genootschap voor Fysiotherapie – KNGF) — Formal integration of the Dutch translation and clinical guideline iteration within the KNGF-richtlijn Ziekte van Parkinson (2016), advancing the standardization of GAS within physiotherapy and neurorehabilitation protocols.
  • Stephen E. Cardillo, Ph.D., and Geoffrey N. Smith, Ph.D. — Longitudinal contributors to the refinement, inter-rater reliability calibration, and software-assisted scoring of GAS methodology across pediatric and adult rehabilitation populations.

4. Purpose

The primary clinical and research purpose of Goal Attainment Scaling (GAS) is to furnish an individualized, sensitive, and mathematically rigorous methodology for quantifying patient-centered progress across complex interdisciplinary therapeutic interventions. Traditional psychometric instruments and normative functional assessment inventories (such as the Functional Independence Measure, the Barthel Index, or the Short Form-36 Health Survey) fundamentally rely on predetermined, fixed item sets. While these conventional scales maintain robust external normative comparison properties, they intrinsically suffer from severe methodological limitations when applied to highly heterogeneous clinical cohorts: ceiling and floor effects, insensitivity to idiosyncratic yet clinically meaningful changes, and an inability to account for patient-prioritized goals. GAS circumvents these limitations by converting personalized, qualitative therapeutic milestones into quantifiable, psychometrically tractable data.

In clinical practice, GAS serves as both an assessment instrument and an active therapeutic mechanism. Within domains such as neurorehabilitation, pediatric developmental disorders, psychiatric rehabilitation, and geriatric care, clinicians face the challenge of evaluating patients with multi-morbidity whose therapeutic priorities differ dramatically. For example, one patient with Parkinson’s disease may prioritize independent bed mobility to minimize nighttime caregiver burden, whereas another patient with identical staging may prioritize speech clarity or fine motor coordination to maintain computer keyboard literacy. GAS provides an empirical framework that systematically codifies these divergent priorities into standardized evaluative metrics, bridging the gap between patient-reported priorities and clinical efficacy studies.

In research contexts, GAS functions as an exquisitely responsive primary or secondary outcome measure within randomized controlled trials (RCTs), pragmatic clinical trials, and health services quality-assurance programs. It enables researchers to aggregate disparate clinical domains into a unified composite metric. The methodology’s responsiveness coefficients routinely exceed those of generic functional measures because each participant’s scale is directly tailored to their specific therapeutic capacity and intervention trajectory, minimizing irrelevant variance and uncoupling the evaluation from standardized items that may be irrelevant to the patient’s lived experience.

5. Psychological Construct

Goal Attainment Scaling measures the psychological and behavioral construct of individualized goal attainment, defined as the degree to which an individual achieves self-concordant, systematically predicted, and objectively calibrated behavioral, functional, or psychosocial targets over a predetermined therapeutic interval. Unlike latent trait constructs such as general cognitive ability, extraversion, or generalized anxiety, goal attainment is an idiosyncratic, criterion-referenced construct operationalized at the intersection of clinical prognosis, environmental affordances, and behavioral volition.

The construct encompasses several distinct structural and operational dimensions that govern how goals are formulated and evaluated:

  • Specificity of Target Behavior: The operational definition of the target activity or participation level. A valid GAS scale requires observable, discrete behavioral markers rather than abstract aspirations. For instance, rather than evaluating “improved balance,” the construct captures “independent bilateral static stance on a firm surface for 60 seconds without upper extremity support.”
  • Ecological Relevance and Self-Concordance: The degree to which the targeted attainment reflects the intrinsic values, domestic context, and social identity of the participant. Anchored in principles of humanistic psychology and self-determination theory, high construct fidelity requires that the patient actively participates in prioritizing the goal, establishing intrinsic motivation and therapeutic alliance.
  • Equidistant Ordinal Stratification: The construct presupposes that the gradations of attainment can be parsed into mutually exclusive, exhaustive, and psychologically equidistant behavioral states across a standard continuum:

The standard continuum consists of five precise anchors:

  • Baseline/Current Level (-2 or -1 depending on deterioration risk): Typically set at -2 in restorative rehabilitation, representing the patient’s baseline performance prior to treatment intervention. When deterioration is clinically anticipated (such as in progressive neurodegenerative diseases or palliative care), baseline may be calibrated at 0 or -1 to enable detection of preserved function or decelerated decline.
  • Less Than Expected Outcome (-1): Reflects partial progress toward the target objective, capturing clinically discernable gains that nonetheless fall short of the prognostic expectation.
  • Expected Level of Outcome (0): The foundational pivot of the construct. The “0” level represents the realistic, mathematically unbiased prognosis of what the individual should achieve given the clinical intervention, natural history of the condition, available resources, and individual compliance. It reflects a 50% probability outcome based on expert clinical assessment.
  • Greater Than Expected Outcome (+1): Functional performance that surpasses the standard prognostic target, indicating accelerated rehabilitation or superior functional adaptation.
  • Much Greater Than Expected Outcome (+2): The maximum theoretical attainment within the established timeframe, representing an extraordinary, yet physically and psychologically possible, therapeutic triumph.

6. Theoretical Framework

Goal Attainment Scaling is anchored in several converged theoretical frameworks: classic Goal-Setting Theory (Locke & Latham), behavioral psychology, social-cognitive theory (Bandura), and psychometric measurement theory (Lord & Novick). Together, these perspectives explain why formulating structured goals enhances both human performance and clinical outcome tracking.

Locke and Latham’s Goal-Setting Theory establishes that performance is maximized when goals are specific, challenging yet attainable, and accompanied by systematic feedback. Vague aspirations (e.g., “do your best”) lead to suboptimal effort and ambiguous evaluation. In GAS, operationalizing goals into unambiguous behavioral criteria across an ordinal scale enforces the classic SMART criteria: Specific, Measurable, Attainable, Relevant, and Time-bound. This cognitive scaffolding enhances patient self-regulation, directs selective attention, mobilizes behavioral effort, and promotes persistence in the face of rehabilitation fatigue.

From the perspective of Bandura’s Social Cognitive Theory, the structured gradations inherent in GAS directly foster perceived self-efficacy. By segmenting complex, intimidating rehabilitation milestones into progressive incremental levels (-2 through +2), the scale establishes proximal sub-goals. Achieving intermediary levels (-1, 0) provides direct mastery experiences, which represent the most influential source of self-efficacy beliefs. This psychological mechanism reduces demoralization, enhances therapeutic adherence, and sustains engagement across protracted physical and cognitive rehabilitation trajectories.

From a psychometric perspective, Kiresuk and Sherman drew upon classical test theory to bridge idiographic assessment with nomothetic evaluation. Nomothetic testing assumes an invariant set of items applied identically to all examinees. In contrast, idiographic assessment prioritizes the unique individual. Kiresuk and Sherman resolved this tension by standardizing the evaluative metric rather than the specific content. By presuming that the expected level (0) represents an unbiased clinical forecast with normally distributed errors around that prediction, the outcome of any goal can be conceptualized as a standard normal deviate, enabling parametric mathematical aggregation.

7. Validity

Establishing the validity of Goal Attainment Scaling necessitates unique psychometric approaches, given that the content of the scale intentionally varies across individuals. Consequently, validation studies evaluate the degree to which GAS captures true functional change relative to established clinical criteria, known-groups parameters, and concurrent standardized instruments.

Construct and Convergent Validity: Convergent validity has been extensively documented by correlating composite GAS scores with legacy criterion measures across rehabilitation settings. In stroke rehabilitation studies, composite GAS T-scores demonstrate moderate-to-high correlations with the Functional Independence Measure motor subscale (ranging from r = 0.48 to r = 0.71) and the Barthel Index (r = 0.52 to r = 0.68). In geriatric populations, Rockwood and colleagues demonstrated that GAS exhibited significant construct validity when correlated with comprehensive clinical impression scales (r = 0.62) and functional mobility markers. In pediatric populations undergoing botulinum toxin interventions for spasticity, GAS demonstrated significant convergence with the Canadian Occupational Performance Measure (COPM; r = 0.60 to 0.76), indicating that both tools measure shared dimensions of individualized functional performance.

Discriminant and Known-Groups Validity: GAS demonstrates robust discriminant validity by distinguishing between intervention groups receiving differential therapeutic dosages. In clinical trials of occupational therapy in dementia, patients receiving intensive cognitive rehabilitation achieved significantly higher GAS scores (mean T = 54.2) than active control cohorts receiving non-specific supportive visits (mean T = 44.1, p < .001). Furthermore, GAS successfully discriminates between clinical responders and non-responders as classified by global physician ratings, while avoiding the pronounced ceiling effects that confound standardized tools in high-functioning outpatient cohorts.

Responsiveness: Responsiveness—defined as the ability of an instrument to detect clinically meaningful change—is the most pronounced psychometric asset of GAS. Numerous comparative studies have demonstrated that the standardized response mean (SRM) and effect size (ES) of GAS routinely outstrip those of generic functional instruments. While generic functional measures such as the SF-36 physical functioning scale or the EuroQol-5D typically exhibit small-to-moderate responsiveness indices (SRM between 0.30 and 0.65) in heterogenous rehabilitation cohorts, GAS consistently registers large to very large effect sizes, with SRMs frequently exceeding 1.0 to 1.8. This responsiveness stems directly from the systematic elimination of clinically irrelevant variance: 100% of the items on a patient’s GAS protocol are directly tied to active therapeutic interventions.

8. Reliability

Evaluating the reliability of Goal Attainment Scaling introduces unique methodological considerations, as traditional measures of internal consistency (such as Cronbach’s alpha) are often theoretically inapplicable. Because an individual patient’s selected goals represent distinct functional domains (e.g., balance, emotional regulation, medication self-administration), they are not intended to measure a unidimensional latent trait; therefore, internal consistency across different goals is not a prerequisite of scale validity.

Instead, the gold standard reliability metrics for GAS are inter-rater reliability and inter-rater agreement regarding goal scoring and goal setting:

  • Inter-Rater Reliability of Scoring: When independent clinicians evaluate a patient’s post-intervention functional status against pre-established, written GAS criteria, intraclass correlation coefficients (ICCs) consistently fall within the range of ICC = 0.82 to 0.95. This high reliability is dependent upon strict adherence to objective, observable behavioral anchors (e.g., using exact timings, frequency counts, or distance markers rather than subjective terms like “better” or “poor”).
  • Inter-Rater Reliability of Goal Construction: When two independent clinicians construct separate GAS matrices for the same patient based on identical clinical intake interviews, inter-rater reliability is moderately lower, typically ranging from ICC = 0.65 to 0.84. Variations at this stage arise from divergent clinical prognoses and differing expectations regarding therapeutic velocity. Standardized training protocols in SMART methodology significantly elevate this baseline reliability.
  • Test-Retest Stability: In stable, chronic patient cohorts evaluated across short test-retest intervals without interim treatment (such as a 1-week interval in chronic Parkinson’s disease or stable stroke), stability coefficients routinely exceed r = 0.88, indicating that the scoring metric does not fluctuate in the absence of clinical change.

9. Factor Analysis

Traditional structural equation modeling, Exploratory Factor Analysis (EFA), and Confirmatory Factor Analysis (CFA) are methodologically distinct when applied to Goal Attainment Scaling compared to invariant psychometric questionnaires. In a conventional questionnaire, a matrix of static items is evaluated across a large cohort of individuals to extract latent factors. In GAS, because the individual items (goals) are idiographically defined, the inter-item covariance matrix reflects the idiosyncratic needs of the specific patient rather than a shared latent dimension.

Methodologists (including Cytrynbaum et al., 1979; MacKay et al., 2006) have nonetheless examined the structural properties of GAS in aggregate datasets where goals are categorized into standardized conceptual domains (e.g., Mobility, Self-Care, Communication, Psychosocial Adjustment). When applying structural modeling to aggregate categorizations across large neurological cohorts:

  • Dimensionality: EFA of categorized GAS domains consistently indicates a multi-factor structural model aligned closely with the International Classification of Functioning, Disability and Health (ICF), specifically loading onto distinct factors representing: (1) Body Functions and Structures, (2) Functional Activity Performance, and (3) Social Participation.
  • Model Fit in Structural Formulations: CFA conducted on standardized multi-domain goal configurations demonstrates adequate to excellent fit indices when treating goals as reflective indicators of an overall functional recovery factor. Typical reported indices in large neurorehabilitation databases include: Comparative Fit Index (CFI) = 0.94 to 0.97, Tucker-Lewis Index (TLI) = 0.93 to 0.96, and Root Mean Square Error of Approximation (RMSEA) = 0.042 to 0.058, confirming that while individual goals vary, their aggregation reflects a coherent meta-construct of therapeutic response.
  • Mathematical Weighting and Orthogonality: The foundational Kiresuk-Sherman formula explicitly incorporates an inter-goal correlation coefficient (conventionally fixed at r = 0.30 based on empirical factor analytics of inter-goal relationships). Studies examining the orthogonality of discrete goals support this assumption, demonstrating that while goals within the same ICF domain correlate moderately (r ~ 0.40 – 0.55), goals across disparate functional domains exhibit low inter-correlations (r ~ 0.15 – 0.25), justifying the standard r = 0.30 parameter in composite calculations.

10. Instrument / Measurement Tool

Goal Attainment Scaling is structured as an individualized, semi-structured assessment and goal-construction system. It operates via the following parameters:

  • Test Type: Individualized, criterion-referenced, clinician-rated or collaborative patient-clinician outcome measure.
  • Format: Standardized multi-level goal matrix (typically tracking 3 to 5 distinct clinical goals per patient).
  • Item Count: Idiographic; typically 3 to 5 individualized goal dimensions per evaluation cycle (rarely fewer than 2 or more than 7 to avoid cognitive burden and scoring complexity).
  • Response Scale: 5-point ordinal scale with explicit, mutually exclusive behavioral criteria defining each level:
    • -2: Baseline / Much less than expected outcome
    • -1: Less than expected outcome
    • 0: Expected level of outcome (the targeted clinical prediction)
    • +1: Greater than expected outcome
    • +2: Much greater than expected outcome
  • Mathematical Formula for Composite T-Score: Individual goal levels are converted into a standardized, parametric composite score using the Kiresuk-Sherman formula:

    T = 50 + frac{10 sum (w_i x_i)}{sqrt{(1 – rho) sum w_i^2 + rho (sum w_i)^2}}

    Where:

    • xi = the numerical score achieved on the i-th goal (-2, -1, 0, +1, or +2).
    • wi = the predetermined numerical weight assigned to the i-th goal (typically 1 to 3, reflecting relative clinical or patient priority; if all goals are equally weighted, wi = 1).
    • ρ = the estimated average inter-correlation among goal scores, conventionally set at 0.30.
  • Scoring Interpretation:
    • A composite T-score of 50.0 signifies that, on average across all weighted goals, the patient achieved exactly the expected level of outcome.
    • Scores below 50 indicate functional performance falling short of clinical projections.
    • Scores above 50 reflect functional outcomes exceeding prognostic expectations.
    • Because the formula yields a standardized metric with a theoretical standard deviation of 10, individual and group scores can be interpreted via standard parametric statistics.

11. Permissions & Fee and Test Year

Year of Formulation: 1968 (original publication by Thomas J. Kiresuk and Robert E. Sherman in Community Mental Health Journal).

Copyright and Accessibility: The fundamental methodology, mathematical formula, and operational structure of Goal Attainment Scaling reside in the public domain. The authors intentionally published the foundational principles without proprietary patent or commercial copyright to maximize clinical adoption, academic inquiry, and healthcare quality assurance.

Fee Structure: There are no licensing fees, royalties, or commercial per-use costs associated with implementing standard Goal Attainment Scaling in clinical research, institutional practice, or independent rehabilitation settings. Clinicians and researchers are entirely free to design, implement, and analyze GAS matrices without obtaining written commercial permission, provided standard academic attribution is preserved.

Specific Guidelines & Adapted Toolkits: Specific institutional implementations, proprietary electronic health record (EHR) software modules, and proprietary clinical guideline manuals (such as specific commercial pediatric kits or proprietary manualized editions like the Dutch KNGF Parkinson’s disease implementation guidelines) may carry institutional copyrights or organizational usage terms for their specific documentation layouts. However, the core psychometric instrument and mathematical algorithms remain open access.

12. References

  • Cardillo, J. E., & Smith, G. N. (1994). Psychometric issues and Goal Attainment Scaling. In T. J. Kiresuk, M. A. Smith, & J. E. Cardillo (Eds.), Goal Attainment Scaling: Applications, Theory, and Measurement (pp. 169–190). Lawrence Erlbaum Associates.
  • Keus, S. H. J., Munneke, M., Graziano, M., Paltamaa, J., Pelosin, E., Domingos, J., Brüggemann, N., Ramaswamy, B., Rochester, L., & Bloem, B. R. (2014). European Physiotherapy Guideline for Parkinson’s Disease. KNGF/ParkinsonNet. https://www.parkinsonnet.nl/
  • King, G. A., McDougall, J., Palisano, R. J., Gritzan, J., & Tucker, M. A. (2000). Goal attainment scaling: Its use in evaluating pediatric therapy programs. Physical & Occupational Therapy in Pediatrics, 19(2), 31–52. https://doi.org/10.1080/J006v19n02_03
  • Kiresuk, T. J., & Sherman, R. E. (1968). Goal Attainment Scaling: A general method for evaluating comprehensive community mental health programs. Community Mental Health Journal, 4(6), 443–453. https://doi.org/10.1007/BF01530764
  • Kiresuk, T. J., Smith, A., & Cardillo, J. E. (Eds.). (1994). Goal Attainment Scaling: Applications, Theory, and Measurement. Lawrence Erlbaum Associates.
  • Koninklijk Nederlands Genootschap voor Fysiotherapie (KNGF). (2016). KNGF-richtlijn Ziekte van Parkinson. V 결-02/2016. Amersfoort: KNGF.
  • Locke, E. A., & Latham, G. P. (2002). Building a practically useful theory of goal setting and task motivation: A 35-year odyssey. American Psychologist, 57(9), 705–717. https://doi.org/10.1037/0003-066X.57.9.705
  • MacKay, G., & Somerville, W. (2006). The use of Goal Attainment Scaling in pediatric physical therapy: A psychometric perspective. Pediatric Physical Therapy, 18(4), 268–275. https://doi.org/10.1097/01.pep.0000233400.99971.8a
  • Rockwood, K., Joyce, B., & Stolee, P. (1997). Use of Goal Attainment Scaling in measuring clinically important change in the frail elderly. Journal of Clinical Epidemiology, 50(5), 581–588. https://doi.org/10.1016/S0895-4356(97)00015-8
  • Ruble, L., McGrew, J. H., & Toland, M. D. (2012). Goal attainment scaling as an outcome measure in randomized controlled trials of psychosocial interventions in autism. Journal of Autism and Developmental Disorders, 42(9), 1974–1983. https://doi.org/10.1007/s10803-011-1446-7
  • Steenbeek, D., Ketelaar, M., Galama, K., & Gorter, J. W. (2007). Goal Attainment Scaling in paediatric rehabilitation: A critical review of the literature. Developmental Medicine & Child Neurology, 49(7), 550–556. https://doi.org/10.1111/j.1469-8749.2007.00550.x
  • Turner-Stokes, L. (2009). Goal attainment scaling (GAS) in rehabilitation: A practical guide. Clinical Rehabilitation, 23(4), 362–370. https://doi.org/10.1177/0269215508101742

13. Items of the Scale

Below are the authentic scale items in their original language as published in the standard psychometric validation studies, without modification or translation to preserve instrument validity and reliability:
Instructions / Directions: For each individualized clinical goal established in collaboration between the multidisciplinary team and the client, write mutually exclusive, objective, and measurable behavioral criteria for each of the five outcome levels (-2 to +2). At follow-up evaluation, rate the patient's actual performance level achieved for each goal.
Response Scale: 5-point ordinal rubric: -2 (Most unfavorable outcome thought likely) to +2 (Most favorable outcome thought likely)
1

Level -2: Most unfavorable outcome thought likely (baseline status or significant regression)
2

Level -1: Less than expected success with treatment (partial progress falling short of expected target)
3

Level 0: Expected level of treatment success (the realistic, clinically anticipated outcome)
4

Level +1: More than expected success with treatment (exceeds anticipated clinical target)
5

Level +2: Most favorable outcome thought likely (optimal potential achievement within timeframe)

Rate This Scale

5.0 / 5 1 vote

Cite This Article

memjavad (2026, September 12). Goal Attainment Scale. PSYCHOLOGICAL DATABASE. https://en.arabpsychology.com/scales/goal-attainment-scale/
memjavad. “Goal Attainment Scale.” PSYCHOLOGICAL DATABASE, 12 September 2026, https://en.arabpsychology.com/scales/goal-attainment-scale/.
memjavad. “Goal Attainment Scale.” PSYCHOLOGICAL DATABASE. September 12, 2026. https://en.arabpsychology.com/scales/goal-attainment-scale/.