Clinical PsychologyPsychometricsPsychotherapy Research

Check List for Clinical Observations

The Check List for Clinical Observations (Forer et al., 1961) is a 44-item observational psychometric tool developed at the Veterans Administration Outpatient Clinic to assess clinical interactions, communicative behaviors, ego defenses, and therapist techniques in psychotherapy.

memjavad
PUBLISHED
Scientifically Reviewed · Dr. Marwa Abd-Alazim · September 28, 2026
Medically & Scientifically Reviewed Verified: September 28, 2026
Dr. Marwa Abd-Alazim Ph.D.
Professor of Psychology • University of Kerbala
Review Criteria & Clinical Standards

This content undergoes rigorous scientific peer-review and medical editorial standards at Arab Psychology Network to ensure clinical accuracy, validity, and compliance with evidence-based guidelines from leading psychological and healthcare authorities (APA / WHO).

Abstract

The Check List for Clinical Observations (Forer et al., 1961) is a specialized 44-item observational rating instrument designed to systematically capture, quantify, and evaluate behavioral and inferential clinical interactions during naturalistic psychotherapy transactions. Developed by a distinguished research team of clinical psychology diplomates at the Veterans Administration Outpatient Clinic in Los Angeles—including Bertram R. Forer, Norman L. Farberow, Herman Feifel, Mortimer M. Meyer, Vita S. Sommers, and Ruth S. Tolman—the instrument operationalizes both directly perceptible therapeutic actions and high-order psychodynamic inferences. The instrument classifies observed phenomena into distinct structural categories: therapist technique and communicative stance, patient expressive behaviors and resistance indicators, primary thematic content domains (e.g., authority, dependency, sexuality, hostility), emotional experiencing and discrete affective states, ego defense mechanisms (such as projection, intellectualization, and isolation), and prognostic/diagnostic clinical appraisals. Each item is scored under a binary observation format denoting presence or absence ("present" vs. "not present") across 50-minute outpatient psychotherapy sessions. Across twelve weekly observation sessions spanning three independent therapist-patient dyads, the checklist underwent systematic empirical refinement to investigate and calibrate interrater consensus among experienced clinicians. Classical interrater reliability investigations yielded phi coefficients ranging from .27 to .70 across individual items, with median phi coefficients settling at .49, .45, and .40 across three successive developmental trial periods. The checklist represents an early milestone in the psychometric measurement of naturalistic psychotherapeutic transactions, illustrating the inherent trade-offs between observational objectivity, clinical inference, and rater concordance in psychiatric and clinical research.

Keywords

Check List for Clinical Observations, clinical interaction, psychotherapy process research, observational checklist, interrater reliability, phi coefficient, clinical inference, psychodynamic mechanisms, Veterans Administration, Bertram Forer, test construction

Authors

The Check List for Clinical Observations was developed in 1961 by a collaborative team of senior clinical psychologists affiliated with the Mental Hygiene Clinic at the Veterans Administration Outpatient Clinic in Los Angeles, California:

  • Bertram R. Forer, Ph.D., ABPP: Widely known for his foundational research in personality assessment, diagnostic validation, and cognitive bias (the Forer effect). Dr. Forer served as a senior clinical psychologist at the Los Angeles VA Outpatient Clinic and held diplomate status with the American Board of Professional Psychology (ABPP).
  • Norman L. Farberow, Ph.D., ABPP: An internationally recognized pioneer in suicidology and co-founder of the Los Angeles Suicide Prevention Center. Dr. Farberow made substantial contributions to clinical crisis intervention, psychodiagnostics, and observational measurement in psychiatric settings.
  • Herman Feifel, Ph.D., ABPP: A seminal figure in modern thanatology and author of The Meaning of Death (1959). Dr. Feifel investigated the existential dimensions of psychotherapy, coping mechanisms, and clinical judgment during long-term outpatient interventions.
  • Mortimer M. Meyer, Ph.D., ABPP: Chief Psychologist at the Veterans Administration Outpatient Clinic in Los Angeles, known for his methodological work in projective techniques, clinical training, and the standardization of psychological evaluations.
  • Vita S. Sommers, Ph.D.: Clinical psychologist and psychotherapist at the Los Angeles VA Clinic, specializing in psychodynamic formulations, ego psychology, and social adaptations of veteran populations.
  • Ruth S. Tolman, Ph.D.: Senior clinical psychologist, researcher, and clinical administrator within the Veterans Administration mental health infrastructure, recognized for her contributions to clinical training, public mental health policy, and psychotherapy evaluation.

Purpose

The primary purpose of the Check List for Clinical Observations is to provide an empirical methodology for recording, classifying, and quantifying the multidimensional events occurring within therapeutic transactions. Prior to the 1960s, naturalistic psychotherapy research was largely dominated by retrospective case summaries, unstandardized clinical notes, and post-session impressions. These subjective narrative modalities presented substantial vulnerabilities to selective memory, theoretical dogmatism, and confirmation bias. Forer and his colleagues conceptualized the checklist as an objective bridge connecting observable clinical phenomenology with systematic psychometric evaluation.

The scale was developed to accomplish several distinct clinical and research objectives:

  • Systematizing Psychotherapeutic Observation: To establish a shared, operationalized behavioral vocabulary among clinical observers viewing ongoing live or recorded psychotherapy hours.
  • Quantifying Levels of Clinical Inference: To differentiate between low-inference observational data (such as measurable silences, motor fidgeting, and explicit conversational interruptions) and high-inference psychodynamic constructs (such as ego defenses, transference configurations, and prognostic viability).
  • Calibrating Interrater Agreement: To evaluate the degree of consensus achievable among expert, board-certified clinical psychologists evaluating identical treatment sessions, thereby illuminating the limits and parameters of clinical consensus.
  • Process Tracking Across Sequential Sessions: To monitor longitudinal therapeutic shifts across serial treatment sessions, including changes in thematic resistance, affective experiencing, defense organization, and therapeutic focus over time.

In clinical research, the instrument serves as an observational framework for comparative process studies, psychotherapy training supervision, and the operational evaluation of dyadic communication. In supervisory contexts, the checklist allows supervisors and trainees to systematically review video or one-way mirror observations, contrasting subjective therapeutic impressions against concrete behavioral markers.

Psychological Construct

The central construct measured by the Check List for Clinical Observations is the Clinical Interaction in Psychotherapy. In the conceptual framework formulated by Forer et al. (1961), therapeutic transactions represent a complex, simultaneous convergence of observable communicative behaviors, relational dynamics, covert psychic conflicts, and structural ego operations. Rather than treating psychotherapy as an undifferentiated dialogue, the instrument stratifies the clinical interaction into several operationalized dimensions:

1. Therapist Behavioral Style and Technical Modality

This dimension assesses the verbal activity level, affective posture, and intervention strategy employed by the clinician. It discriminates whether the therapist maintains an active verbal presence versus a passive listener stance, displays clinical comfort and ease, and relies predominantly on reflective, interpretive, or supportive therapeutic interventions.

2. Temporal and Communicative Session Structure

This domain captures the thematic coherence and conversational pacing of the therapeutic encounter. It measures whether the hour centers upon a focal psychological problem or is characterized by patient topic-rambling. It also records behavioral silences (defined operationally as non-verbal pauses lasting 30 seconds or longer) and identifies whether silences are systematically broken by the therapist or the client.

3. Manifestations of Clinical Resistance

Resistance is evaluated across both behavioral markers and communicative patterns. The construct operationalizes verbal dispute of therapist interpretations, conversational interruptions, halting speech patterns with hesitations and incomplete utterances, defensive over-generalization, and the evasive utilization of abstract psychological terminology.

4. Thematic Problem Domains

The instrument categorizes the manifest psychological content brought into the therapeutic session. Observers evaluate whether active work focuses on authority dynamics, psychosexual functioning, dependency conflicts, occupational stressors, aggressive and hostile impulses, emotional self-regulation, somatic or psychiatric symptoms, or interpersonal relationships.

5. Expressive Motor Behavior and Productivity

This category assesses motoric non-verbal cues (such as physical gestures, restlessness, fidgeting, and flat monotone speech) as well as the quantitative volume and substantive depth of patient verbal material. Substantive depth is operationalized as the capacity to meaningfully link present psychological functioning to past developmental antecedents or the presentation of core material accompanied by intense emotional resonance.

6. Affective Experiencing and Emotional State

This dimension measures the overt experiencing of emotional states and discriminates specific qualitative affects displayed during the hour, including anger, fear, sadness, anxiety, and warmth.

7. Characterological and Ego Capacities

The checklist assesses ego adaptability and psychological mindedness, including behavioral rigidity during the session, relational capacity, potential for clinical insight, degree of self-criticism, and the conscious affective valence of the working relationship.

8. Defense Mechanisms and Prognostic Appraisal

At the highest level of clinical inference, observers evaluate classic psychoanalytic ego defenses—such as projection, repression, rationalization, denial/avoidance, intellectualization, isolation of affect, reaction formation, and displacement/conversion. Finally, clinicians synthesize these observations into longitudinal prognostic judgments regarding treatment completion, estimated duration of care, and formal diagnostic impression.

Theoretical Framework

The Check List for Clinical Observations is anchored within mid-twentieth-century ego psychology, psychodynamic theory, and empirical communication science. During the 1950s and 1960s, American clinical psychology—spurred heavily by Veterans Administration training centers—sought to reconcile psychoanalytic theories of internal conflict with the scientific demands of observable, empirical psychometrics. This integration drew upon foundational paradigms advanced by Sigmund Freud, Anna Freud, and Heinz Hartmann, alongside emerging interactional process theories.

Ego Psychology and Defense Operationalization

Anna Freud’s landmark work, The Ego and the Mechanisms of Defence (1936), positioned ego defenses as the primary observable manifestations of unconscious psychic adaptation. Forer et al. operationalized these theoretical constructs into discrete observational criteria. Rather than conceptualizing defense mechanisms solely as unobservable intrapsychic events, the authors defined them through manifest linguistic and interactional behaviors: intellectualization as abstract generalization, isolation as the dissociation of affect from verbal narrative, and projection as the overt externalization of unacceptable personal impulses.

The Inference Gradient in Clinical Judgment

A central theoretical premise of Forer and colleagues’ work is the continuum of clinical inference. Observational phenomena in clinical settings do not reside at a single epistemological level. Instead, they occupy a spectrum ranging from:

  1. Directly Perceptible Physical Events (Low Inference): Measuring physical silence duration (e.g., >30 seconds), physical fidgeting, or verbal interruption frequencies.
  2. Contextual Interactional Assessments (Moderate Inference): Judging whether the session maintained thematic focus, evaluating patient resistance, or determining the presence of warmth.
  3. Structural Dynamic Formulations (High Inference): Imputing unconscious defense mechanisms, evaluating underlying transference configurations, and forecasting long-term therapeutic prognosis.

By mapping these distinct strata within a single rating protocol, the theoretical framework of the checklist allows researchers to examine how observational consensus degrades or holds stable as raters move from concrete behavioral counts to abstract clinical deductions.

Validity

In the original developmental investigation published by Forer, Farberow, Feifel, Meyer, Sommers, and Tolman (1961), formal statistical validation centered primarily on content validity, expert consensus validity, and ecological representativeness rather than classical criterion-oriented or psychometric construct validation algorithms.

Content and Face Validity

The items of the checklist were generated through iterative clinical seminars conducted by six senior diplomates of the American Board of Examiners in Professional Psychology (ABEPP/ABPP). Over repeated observational cycles, items were drafted, field-tested, and refined to ensure direct correspondence with the clinically salient events of outpatient psychotherapy. The resulting item pool captures the full operational scope of individual psychotherapy sessions, providing robust face and content validity for adult outpatient treatment evaluation.

Iterative Empirical Refinement

The instrument underwent formal structural calibration across 12 consecutive weekly psychotherapy sessions involving three distinct therapist-patient pairs. After every three sessions, the research team analyzed item performance, observer discrepancies, and linguistic ambiguities. Clarification criteria were added to items demonstrating ambiguous boundaries (such as adding behavioral qualifiers to "resistance" or defining "depth of material" as connecting past experiences to present functioning). Two items—Item 19 ("Relationships with people") and Item 32 ("Patient seems self-critical")—were formally introduced into the scale following the initial observational block to rectify observed content omissions.

Construct and Discriminant Considerations

Although the original monograph reported no formal correlational validity matrices with outside criterion batteries, the authors observed clear discriminant patterns across differing levels of inference. Low-inference behavioral items demonstrated distinct functional independence from high-inference psychodynamic constructs. High-inference items (such as specific ego defense mechanisms and diagnostic categorizations) exhibited lower inter-judge convergence, reflecting the psychometric vulnerability of abstract interpretive constructs to idiosyncratic theoretical predilections.

Reliability

The primary psychometric evaluation of the Check List for Clinical Observations centered upon interrater reliability and percentage of observer agreement across independent clinical judges.

Interrater Agreement and Phi Coefficients

Because items were scored in a dichotomous binary format ("present" vs. "not present"), the investigators utilized the phi coefficient ($phi$), a statistical measure of association for two binary variables equivalent to the Pearson product-moment correlation coefficient computed for dichotomous categories. Across the three successive observation periods, the range of phi coefficients spanned from .27 to .70 across the evaluated items.

Stability Across Observation Periods

Across the three sequential four-week observational blocks, median phi coefficients showed slight downward shifts:

  • Observation Period 1 (Sessions 1–3): Median $phi = .49$
  • Observation Period 2 (Sessions 4–6): Median $phi = .45$
  • Observation Period 3 (Sessions 7–12, including revised scale): Median $phi = .40$

The investigators determined that overall item agreement ranged between 70% and 90% across raw scoring protocols. However, because binary base rates for specific rare behaviors (e.g., specific isolated defense mechanisms or rare affects) were skewed, chance-corrected association coefficients settled within moderate ranges. Low-inference items specifying discrete time intervals (such as silences lasting 30 seconds or more) or physical actions (such as motor gestures and postural restlessness) consistently demonstrated the highest reliability coefficients ($phi > .60$). Conversely, high-inference items requiring qualitative clinical interpretation (such as reaction formation or unconscious resistance) yielded lower reliability indices ($phi < .35$).

Factor Analysis

No formal exploratory factor analysis (EFA) or confirmatory factor analysis (CFA) was performed in the original 1961 publication. The statistical conventions of the late 1950s and early 1960s, combined with the sample size of 12 clinical sessions across three dyads, precluded the computation of large-scale correlation matrices necessary for stable factor extraction.

Instead, the 44 items were categorized based on clinical and theoretical relevance:

  • Therapist Posture and Strategy (Items 1–5): Assessing therapist activity level, comfort, and therapeutic approach.
  • Session Focus and Temporal Dynamics (Items 6–8): Capturing thematic focus, silences, and who breaks pauses.
  • Resistance Manifestations (Item 9, Sub-items 9a–9e): Capturing verbal and stylistic non-cooperation.
  • Expressive Patient Behaviors (Items 10–11, 20–22): Measuring speech tone, motor agitation, volume, and material depth.
  • Thematic Problem Areas (Items 12–19): Content analysis of core life conflicts.
  • Affective Spectrum (Items 23–28): General affect and specific emotion ratings.
  • Ego Functioning and Prognosis (Items 29–33, 41–44): Rigidity, insight, rapport, and longitudinal prognosis.
  • Defense Mechanisms (Items 34–40): Psychodynamic ego adaptations.

Subsequent psychotherapy process research (e.g., early structural equation modeling and factor analytic investigations of session evaluation inventories) has consistently validated these conceptual clusters as empirical domains of therapeutic process ratings.

Instrument / Measurement Tool

The Check List for Clinical Observations is structured as an observer-rated checklist intended for use during or immediately following clinical psychotherapy sessions. Below are the operational specifications of the measurement tool:

  • Test Type: Observer-rated clinical process checklist.
  • Administration Method: Live observation (via one-way mirror) or systematic video/audio recording review of 50-minute clinical sessions.
  • Target Population: Adult psychotherapy clients (18 years of age and older) and their treating clinicians across outpatient or inpatient psychiatric contexts.
  • Target Raters: Trained mental health professionals, clinical psychologists, psychiatric residents, or advanced clinical trainees.
  • Item Count: 44 primary items (incorporating supplementary clinical sub-items and operational clarifications).
  • Response Format: Binary dichotomous judgment: Present versus Not Present for each item during the evaluated session.
  • Scoring and Quantification:
    • Individual items are evaluated independently for frequency or occurrence during the 50-minute session.
    • Sub-items (e.g., Item 8a/b, Items 9a–9e, Items 12–19, Items 24–28) denote specific qualitative manifestations within a broader clinical domain.
    • Composite sub-scores can be derived by summing present indicators across specific dimensions (e.g., Total Resistance Markers, Affective Range, Defense Repertoire).
    • Interrater consensus is determined via concordance percentage or binary correlation coefficients (phi coefficient).

Permissions & Fee and Test Year

The Check List for Clinical Observations was originally published in 1961 in the Journal of Consulting Psychology by Bertram R. Forer, Norman L. Farberow, Herman Feifel, Mortimer M. Meyer, Vita S. Sommers, and Ruth S. Tolman. As research conducted primarily under the auspices of the Veterans Administration Outpatient Clinic and published in standard academic literature over six decades ago, the scale is widely accessible for non-commercial scholarly and research purposes. Formal publication copyright was historically held by the American Psychological Association (APA). Researchers seeking to reproduce the instrument in commercial assessment batteries or published clinical handbooks should consult the permissions guidelines of the APA or the original publication source (DOI: 10.1037/h0044887). No testing fees are required for academic research applications.

References

  • Feifel, H. (Ed.). (1959). The meaning of death. McGraw-Hill. https://doi.org/10.1037/11189-000
  • Forer, B. R. (1949). The fallacy of personal validation: A classroom demonstration of gullibility. The Journal of Abnormal and Social Psychology, 44(1), 118–123. https://doi.org/10.1037/h0059240
  • Forer, B. R., Farberow, N. L., Feifel, H., Meyer, M. M., Sommers, V. S., & Tolman, R. S. (1961). Clinical perception of the therapeutic transaction. Journal of Consulting Psychology, 25(2), 93–101. https://doi.org/10.1037/h0044887
  • Freud, A. (1936). The ego and the mechanisms of defence. International Universities Press.
  • Orlinsky, D. E., & Howard, K. I. (1975). Varieties of psychotherapeutic experience: Multivariate analyses of patients’ and therapists’ reports. Teachers College Press.
  • Shneidman, E. S., & Farberow, N. L. (Eds.). (1957). Clues to suicide. McGraw-Hill.

Items of the Scale

Below are the authentic scale items in their original language as published in the standard psychometric validation studies, without modification or translation to preserve instrument validity and reliability:

Response Scale: Each item is evaluated as Present or Not Present.

  1. Therapist was primarily active. (Active: verbal activity)a
  2. Therapist was primarily comfortable. (Comfortable: has sense of ease with the patient. Frame of reference should be our concept of comfort for all therapists.)
  3. Reflective
  4. Interpretive
  5. Supportive
  6. The session seemed to be focused on a problem.
  7. There were silences. (Silence: period when no one is talking, 30 seconds or more)
  8. If there were silences, they were usually broken by:
  9. The hour was characterized by resistance.
  10. Patient spoke in a monotone.
  11. Patient was fidgety or restless.
  12. Authority
  13. Sex
  14. Dependency
  15. Work
  16. Hostility
  17. Emotional control
  18. Symptoms
  19. Relationships with peopleb
  20. Patient used gestures. (More than once)
  21. Patient brought up a lot of material. (Material: variety or elaboration of content)
  22. The material brought up was deep. (Deep: (a) patient meaningfully relates something that is happening in present to what has happened in past [content], or (b) produces something with great affect)
  23. Patient was experiencing affect.
  24. Anger
  25. Fear
  26. Sadness
  27. Anxiety
  28. Warmth
  29. Patient seems to be rigid. (This applies to fluidity and spontaneity during the hour, whether material is censored or uncensored, etc. Not a judgment of character structure.)
  30. Patient seems to show ability to form close relationships.
  31. Patient seems capable of insight.
  32. Patient seems self-critical.b
  33. (Patient’s relationship to the therapist during the hour seemed to be one of positive feelings or rapport. This refers to conscious feelings.)
  34. Projection (Attribution to other person of motives unacceptable to oneself).
  35. Repression (Rationalization: logical excuse or justification of feelings or behavior).
  36. Avoidance (Denial: avowed nonperception of reality situation, internal or external).
  37. Intellectualization (Explanation of one’s feelings or behavior in terms of general or theoretical principles or abstract concepts).
  38. Isolation (Separation of feeling and idea).
  39. Reaction formation (Turning into the opposite, internal).
  40. Conversion (Displacement: shift of feeling from one object or person to another).
  41. Patient will complete therapy.
  42. Patient will need long-term therapy, 1 ½ years or more.
  43. Diagnostic impression.
★

Rate This Scale

5.0 / 5 • 1 vote

Cite This Article

memjavad (2026, September 28). Check List for Clinical Observations. PSYCHOLOGICAL DATABASE. https://en.arabpsychology.com/scales/check-list-for-clinical-observations/
memjavad. “Check List for Clinical Observations.” PSYCHOLOGICAL DATABASE, 28 September 2026, https://en.arabpsychology.com/scales/check-list-for-clinical-observations/.
memjavad. “Check List for Clinical Observations.” PSYCHOLOGICAL DATABASE. September 28, 2026. https://en.arabpsychology.com/scales/check-list-for-clinical-observations/.