1. Abstract
The quantification of unobservable psychological, behavioral, and epidemiological constructs is a foundational endeavor across the social, behavioral, and clinical health sciences. Because latent variables—such as psychological distress, perceived stigma, food insecurity, and health-related self-efficacy—cannot be measured directly via physical instrumentation, investigators rely on psychometric scales composed of observable multi-item indicators. However, the development of psychometrically sound measurement tools is an intricate, multi-stage scientific endeavor that requires methodological rigor across theoretical, statistical, and operational dimensions. The framework articulated by Godfred O. Boateng, Torsten B. Neilands, Edward A. Frongillo, Hugo R. Melgar-Quiñonez, and Sera L. Young (2018) provides a comprehensive, nine-step primer systematically organized into three distinct developmental phases: Item Development, Scale Development, and Scale Evaluation.
In Phase I (Item Development), researchers identify the target construct’s theoretical boundaries, generate a preliminary item pool through deductive and inductive modalities, and evaluate content validity using expert adjudication and cognitive pre-testing interviews. In Phase II (Scale Development), the instrument is pre-tested on target populations, subjected to item reduction strategies (including item discrimination metrics, inter-item correlations, and classical test theory thresholds), and analyzed via factor extraction techniques—primarily exploratory factor analysis (EFA) or Rasch and Item Response Theory (IRT) models—to ascertain structural dimensionality. In Phase III (Scale Evaluation), the latent dimensionality is independently replicated through confirmatory factor analysis (CFA), structural equation modeling, and measurement invariance evaluations across diverse demographic strata. Finally, psychometric adequacy is determined via composite reliability indices (Cronbach’s alpha, McDonald’s omega, ordinal theta) and empirical validity appraisals, encompassing convergent, discriminant, and criterion-related validity testing. By delineating a systematic, evidence-based roadmap, this primer safeguards against widespread measurement error, curbs construct contamination, and elevates the reproducibility and clinical utility of psychometric scales worldwide.
2. Keywords
Scale Development, Psychometrics, Construct Validity, Measurement Invariance, Exploratory Factor Analysis, Confirmatory Factor Analysis, Classical Test Theory, Item Response Theory, Reliability, Latent Constructs
3. Authors
The foundational primer on scale development and validation was authored by a multidisciplinary team of methodological experts, epidemiologists, and behavioral scientists:
- Godfred O. Boateng, Ph.D. — Department of Anthropology and Institute for Policy Research, Northwestern University, Evanston, IL, United States; currently affiliated with the College of Health Sciences, University of Texas at Arlington, Arlington, TX, United States. Corresponding Email: [email protected]
- Torsten B. Neilands, Ph.D. — Center for AIDS Prevention Studies (CAPS), Department of Medicine, University of California, San Francisco, San Francisco, CA, United States.
- Edward A. Frongillo, Ph.D. — Department of Health Promotion, Education, and Behavior, Arnold School of Public Health, University of South Carolina, Columbia, SC, United States.
- Hugo R. Melgar-Quiñonez, M.D., Ph.D. — Institute for Global Food Security, School of Human Nutrition, McGill University, Montreal, QC, Canada.
- Sera L. Young, Ph.D. — Department of Anthropology and Institute for Policy Research, Northwestern University, Evanston, IL, United States.
4. Purpose
The primary purpose of the methodological framework established by Boateng et al. (2018) is to demystify, systematize, and standardize the complex processes governing the construction, refinement, and psychometric evaluation of self-report instruments. Within behavioral medicine, clinical psychology, public health, and social epidemiology, investigators routinely encounter emergent, nuanced phenomena—such as household water insecurity, vaccine hesitancy, patient empowerment, and intersectional stigma—that lack standardized measurement tools. In the absence of formal graduate-level training in advanced psychometric theory, researchers frequently resort to ad hoc, unvalidated questionnaires or execute fragmented analytical procedures that fail to establish true measurement integrity. Such poorly calibrated instruments inject substantial systematic and random measurement error into empirical research, attenuation of effect sizes, false-negative discoveries, and compromised patient care decisions.
To counteract the proliferation of subpar measurement tools, this framework operationalizes psychometrics into a structured, step-by-step roadmap. Clinically, the protocol ensures that patient-reported outcome measures (PROMs) capture genuine physiological, behavioral, and affective changes rather than semantic ambiguities or differential item functioning. Methodologically, it bridges the historical divide between traditional Classical Test Theory (CTT) and contemporary latent variable modeling paradigms, such as Item Response Theory (IRT) and structural equation modeling (SEM). By codifying explicit thresholds for item retention, structural evaluation, and cross-cultural invariance, the primer serves as an authoritative guide for investigators seeking to build novel diagnostic tools from theoretical inception or critically evaluate and adapt existing global health instruments.
5. Psychological Construct
In psychometric measurement theory, a latent construct is defined as an unobservable theoretical abstraction representing an individual’s cognitive, affective, behavioral, or experiential state. Because latent variables cannot be directly registered via physical sensory apparatus, they must be inferred through a constellation of observable indicators (i.e., manifest scale items). The Boateng et al. framework conceptualizes construct formulation through three essential methodological phases containing nine sequential steps:
Phase I: Item Development
- Step 1: Domain Identification and Item Generation. The construct must be delineated through rigorous theoretical boundaries to establish construct purity and prevent domain underrepresentation or construct contamination. Item pools are developed using deductive strategies (comprehensive literature reviews, clinical diagnostic manuals, existing theoretical frameworks) and inductive strategies (qualitative semi-structured interviews, focus group discussions, cognitive mapping with the target population). This dual approach ensures that the resulting pool reflects both professional domain expertise and lived participant experiences.
- Step 2: Content Validity Assessment. Content validity pertains to the degree to which an instrument’s items completely and accurately reflect the targeted theoretical universe. Subject-matter experts (SMEs) systematically quantify item relevance and clarity using established psychometric metrics, such as Lawshe’s Content Validity Ratio (CVR) and Lynn’s Content Validity Index (CVI). Items failing to achieve pre-established quantitative thresholds are revised or eliminated prior to empirical field testing.
- Step 3: Cognitive Pre-Testing. Before quantitative deployment, drafted items are administered to members of the target demographic using qualitative cognitive interviewing techniques (e.g., “think-aloud” protocols and verbal probing). This step diagnoses hidden semantic discrepancies, sociocultural misinterpretations, recall burdens, or sensitive phrasing that could distort respondent comprehension and introduce systematic measurement error.
Phase II: Scale Development
- Step 4: Scale Administration and Sampling. The refined item pool is administered to a sufficiently powered, representative sample of respondents. The framework emphasizes adequate respondent-to-item ratios (ranging from 5:1 to 10:1 or minimum sample sizes exceeding N = 200 to 300) to ensure parameter stability in multivariate reduction.
- Step 5: Item Reduction. Empirical evaluation of individual item functioning is conducted using Classical Test Theory metrics. Analysts inspect response frequencies to identify floor and ceiling effects, evaluate inter-item correlation matrices (retaining items within the optimal bivariate correlation window of r = 0.20 to 0.70), and compute item-total correlations (flagging items exhibiting corrected correlations below 0.30). Non-functioning distractors and items characterized by high missingness or severe skewness are systematically pruned.
- Step 6: Extraction of Latent Factors. To reveal the intrinsic dimensional architecture of the item pool, the dataset is subjected to exploratory structural analysis. Analysts employ exploratory factor analysis (EFA) with appropriate extraction methods (e.g., principal axis factoring or robust maximum likelihood) and oblique factor rotations (e.g., Promax, Oblimin) to reflect theoretical inter-factor correlations, or implement Rasch/IRT models for unidimensional scaling.
Phase III: Scale Evaluation
- Step 7: Tests of Dimensionality and Model Verification. The empirical factor structure derived in Step 6 is subsequently verified using Confirmatory Factor Analysis (CFA) or Exploratory Structural Equation Modeling (ESEM) on an independent, hold-out validation sample. Structural integrity is confirmed by evaluating global goodness-of-fit statistics alongside local parameter estimates.
- Step 8: Reliability Assessment. The scale’s capacity to yield dependable, internally consistent, and temporally stable observations is quantified using modern composite reliability estimators (Cronbach’s alpha, McDonald’s omega coefficient, ordinal theta) alongside longitudinal test-retest reliability designs evaluated via intraclass correlation coefficients (ICC).
- Step 9: Validity Testing and Nomological Placement. The final instrument is embedded within a nomological network of related and unrelated empirical measures. The primer guides investigators through convergent validity, discriminant validity (e.g., Campbell and Fiske’s Multitrait-Multimethod Matrix, Fornell-Larcker criteria), and criterion validity (predictive and concurrent associations with external clinical endpoints).
6. Theoretical Framework
The Boateng et al. primer is grounded in the convergence of two dominant measurement frameworks: Classical Test Theory (CTT) and Modern Latent Variable Modeling (comprising Item Response Theory and Structural Equation Modeling).
Classical Test Theory, historically conceptualized by Spearman and formalized by Novick, Lord, and Nunnally, operates on the foundational assumption that an observed test score ($X$) is a linear function of an unobservable true score ($T$) and an unsystematic random error component ($E$):
X = T + E
Under CTT axioms, random errors are assumed to have an expected value of zero, be uncorrelated with true scores, and possess mutually uncorrelated variance across parallel test administrations. CTT provides the mathematical rationale for multi-item aggregations: by summing or averaging multiple observed indicators tapping the same latent domain, idiosyncratic item-level random measurement error cancels out, causing composite scores to approach the respondent’s true score level with greater precision.
Concurrently, the primer integrates Modern Latent Variable Theory, which shifts analytical focus from total test scores to the probabilistic interaction between an individual’s underlying trait level ($ heta$) and the mathematical properties of individual items. Within this paradigm, Item Response Theory (IRT) and Rasch modeling provide item difficulty, item discrimination, and item information curves that remain theoretically sample-independent. Furthermore, the structural framework incorporates factor analytic and structural equation modeling principles articulated by Jöreskog, Thurstone, and Bentler. Latent variables are modeled as common factors responsible for the shared variance among observed indicators, separating common construct variance from unique, item-specific variance and measurement error. Consequently, the primer synthesizes CTT’s practical operational utility with the analytical sophistication of confirmatory latent variable structures.
7. Validity
Within the Boateng et al. architecture, validity is conceptualized not as a static property inherent to an instrument, but as an evolving, evidence-based argument supporting the interpretations and clinical inferences drawn from scale scores, aligning with Messick’s unified validity framework. The primer delineates explicit empirical and judgmental procedures to document validity across multiple strata:
Content and Face Validity
Content validity is established early in Phase I through structured quantitative protocols. Subject matter panels evaluate each proposed item using Likert-type ordinal ratings regarding relevance, clarity, and essentiality. Lawshe’s Content Validity Ratio (CVR) is calculated using the formula:
CVR = (ne – N/2) / (N/2)
where ne represents the number of panelists rating the item as “essential,” and N is the total panel size. Items failing to achieve statistical significance based on minimum Lawshe cutoffs (e.g., CVR ≥ 0.62 for 10 panelists) are pruned. Concurrently, Lynn’s Item-Level Content Validity Index (I-CVI) requires an agreement proportion ≥ 0.78 across multiple experts, while the Scale-Level Content Validity Index (S-CVI/Ave) must achieve ≥ 0.90 to confirm domain comprehensiveness.
Construct Validity: Convergent and Discriminant Validation
Construct validity evaluates whether scale scores behave in theoretical alignment with established scientific paradigms. Boateng et al. mandate empirical verification through:
- Convergent Validity: Demonstrated when the newly developed scale correlates moderately to strongly (typically Pearson r ≥ 0.50, p < 0.001) with pre-existing, validated instruments measuring identical or theoretically overlapping constructs. In confirmatory structural models, convergent validity is substantiated when standardized factor loadings exceed 0.50 (ideally ≥ 0.70) and the Average Variance Extracted (AVE) surpasses 0.50, indicating that the latent construct explains over half of its indicators’ variance.
- Discriminant Validity: Confirmed when the scale exhibits low to negligible correlations with instruments designed to measure conceptually distinct, unrelated phenomena (e.g., separating water insecurity from generalized psychological depression). Following the rigorous Fornell-Larcker criterion, discriminant validity is achieved when the square root of the AVE for a given latent factor exceeds its bivariate correlation coefficients with any other latent construct in the structural model.
Criterion and Predictive Validity
Criterion-related validity reflects how accurately scale scores predict an external behavioral, physiological, or diagnostic benchmark. In concurrent validity testing, the instrument is correlated with an established gold standard measured concurrently. In predictive validity testing, longitudinal designs are employed to determine whether baseline scale scores successfully forecast prospective outcomes (e.g., baseline water insecurity predicting subsequent child diarrheal incidence or linear growth faltering at 12-month follow-up).
8. Reliability
Reliability represents the ratio of true score variance to observed score variance, quantifying the extent to which an instrument produces consistent, reproducible measurements free from unsystematic measurement error. Boateng et al. emphasize that while high reliability is a prerequisite for validity, reliability alone does not guarantee construct validity. The primer outlines rigorous methodological criteria for evaluating multiple reliability facets:
Internal Consistency Reliability
Internal consistency appraises the degree of inter-relatedness among items composing a subscale. Historically, Cronbach’s alpha ($lpha$) has been the standard metric:
α = [k / (k – 1)] × [1 – (∑ σi2) / σt2]
The primer highlights that alpha relies on stringent, frequently violated statistical assumptions: essential tau-equivalence (equal factor loadings across all items) and uncorrelated item error variances. When tau-equivalence is violated, Cronbach’s alpha substantially underestimates true scale reliability. Consequently, Boateng et al. advise computing McDonald’s omega ($\omega$) (composite reliability) derived from confirmatory factor analytic loadings:
ω = (∑ λi)2 / [(∑ λi)2 + ∑ θii]
where $\lambda_i$ denotes standardized factor loadings and $ heta_{ii}$ represents unique item error variances. In addition, when items utilize ordinal Likert formats characterized by skewness, the primer recommends calculating ordinal alpha or ordinal theta based on polychoric correlation matrices. Benchmark thresholds recommended across the literature dictate:
- α / ω ≥ 0.70: Acceptable for early exploratory research.
- α / ω ≥ 0.80: Desirable for robust clinical and empirical investigations.
- α / ω > 0.90 to 0.95: Mandatory for individual-level diagnostic decision-making, though values exceeding 0.95 may signal redundant item content.
Temporal Stability: Test-Retest Reliability
For constructs conceptualized as stable psychological traits (e.g., generalized self-efficacy or locus of control), the primer mandates the evaluation of test-retest reliability across repeated administrations. The time interval must be carefully calibrated—typically between 2 and 4 weeks—to prevent memory recall artifacts while minimizing genuine developmental trait maturation. The framework cautions against using simple bivariate Pearson correlations, which only measure linear association while ignoring systematic mean shifts; instead, researchers are directed to employ two-way mixed-effects Intraclass Correlation Coefficients (ICC) modeling absolute agreement, where an ICC ≥ 0.75 reflects satisfactory temporal stability.
9. Factor Analysis
Factor analysis represents the mathematical core of scale development, enabling researchers to explore and confirm the dimensional architecture of an item pool. The framework articulates a dual-stage analytical workflow dividing exploratory and confirmatory strategies across distinct, independent participant samples.
Exploratory Factor Analysis (EFA)
EFA is executed during Phase II to determine the number of latent factors underlying the preliminary item set. Rather than relying on outdated rules of thumb like the Kaiser criterion (eigenvalues > 1.0, which frequently over-factors), the primer advises triangulating multiple criteria:
- Parallel Analysis (Horn’s Method): Comparing empirical sample eigenvalues against the 95th percentile of eigenvalues generated from synthetic random correlation matrices. Factors are retained only if empirical eigenvalues exceed simulated noise thresholds.
- Cattell’s Scree Plot: Visual inspection of the eigenvalue descent curve to identify the inflection point prior to the scree eluvium.
- Minimum Average Partial (MAP) Test: Evaluating partial correlations to retain components that minimize remaining shared variance.
For factor extraction, robust maximum likelihood (ML) or principal axis factoring (PAF) is recommended over principal components analysis (PCA), as PCA assumes zero measurement error and inflates factor loadings. Oblique rotations (e.g., Promax, Direct Quartimin) should be chosen over orthogonal rotations (e.g., Varimax) because latent behavioral and psychological constructs are rarely completely independent ($r = 0$). Items displaying standardized factor loadings < 0.30 to 0.40, or complex cross-loadings differing by less than 0.15 across factors, are systematically flagged for deletion.
Confirmatory Factor Analysis (CFA) and Fit Evaluation
In Phase III, CFA is conducted on an independent validation sample to formally test whether the theoretical factor model accurately reproduces the observed covariance matrix. Model adequacy is evaluated using standardized global fit indices:
- Chi-Square to Degrees of Freedom Ratio ($\chi^2/df$): Values ≤ 3.0 (or ≤ 2.0 in conservative paradigms) suggest acceptable fit, recognizing that absolute $\chi^2$ tests are hypersensitive to large sample sizes.
- Comparative Fit Index (CFI) and Tucker-Lewis Index (TLI): Thresholds ≥ 0.90 indicate adequate fit, while values ≥ 0.95 reflect good model specification.
- Root Mean Square Error of Approximation (RMSEA): Values ≤ 0.06 indicate close fit, while values between 0.06 and 0.08 reflect reasonable approximation error (accompanied by a 90% confidence interval upper limit < 0.08).
- Standardized Root Mean Square Residual (SRMR): Values ≤ 0.08 indicate acceptable residual covariance fit.
Measurement Invariance
To establish that a newly developed scale functions equivalently across diverse cultural, demographic, or clinical subpopulations, the framework prescribes multigroup confirmatory factor analysis (MGCFA). Investigators sequentially test hierarchical levels of invariance: Configural Invariance (identical factor structure across groups), Metric Invariance (equal factor loadings, validating cross-group structural equivalence), Scalar Invariance (equal item intercepts, permitting direct comparison of latent factor means), and Strict Invariance (equal unique item residual variances). Degradation of fit ($Delta ext{CFI} > 0.010$ or $Delta ext{RMSEA} > 0.015$) indicates differential item functioning, requiring partial invariance modeling or cross-cultural recalibration.
10. Instrument / Measurement Tool
The operational characteristics of measurement instruments developed under the Boateng et al. methodological framework adhere to standardized psychometric conventions:
- Test Type: Multi-item, structured psychometric measurement tool (predominantly designed for self-report surveys, patient-reported outcome measures [PROMs], or structured clinical/enumerator interviews).
- Target Population: Adaptable across general community adult populations, specialized clinical cohorts, and diverse cross-cultural settings.
- Structural Modality: Unidimensional or multidimensional latent variable scales, configured into discrete thematic subscales reflecting target theoretical domains.
- Typical Item Volume: Final validated scales typically retain between 10 and 30 items, optimizing psychometric coverage while minimizing respondent fatigue.
- Response Formats: Multi-point Likert scales (e.g., 4-point, 5-point, or 7-point scales anchored from 1 = Strongly Disagree to 5 = Strongly Agree, or frequency anchors from 0 = Never to 4 = Always). Binary formats (0 = No, 1 = Yes) are evaluated using Rasch/IRT dichotomous models.
- Standard Administration Procedures: Paper-and-pencil questionnaires, Computer-Assisted Personal Interviewing (CAPI), or web-based informatics platforms (e.g., REDCap, Qualtrics). Typical completion duration ranges from 5 to 15 minutes.
- Scoring and Transformation Algorithms: Negatively phrased/reverse-keyed items are inverted prior to composite aggregation. Scores are computed via linear subscale summation, mean composite scoring, or standardized factor/IRT latent trait score ($ heta$) estimations. Total scores are frequently transformed onto standardized metric scales (e.g., 0 to 100) or categorized according to empirically validated clinical cut-points.
11. Permissions & Fee and Test Year
The landmark methodological primer “Best Practices for Developing and Validating Scales for Health, Social, and Behavioral Research” was formally published in 2018 in Frontiers in Public Health. The publication is distributed under the terms of the Creative Commons Attribution License (CC BY 4.0). As an open-access methodological treatise, the nine-step framework, developmental guidelines, and conceptual workflows are freely accessible to researchers, clinicians, academic faculties, and global health practitioners worldwide without licensing fees or financial encumbrances.
Investigators who develop specific, standalone psychometric instruments (e.g., the Household Water Insecurity Experiences [HWISE] Scale, Coping Self-Efficacy Scales, or Breastfeeding Support Instruments) utilizing this primer retain separate proprietary or open-access copyrights over their distinct questionnaire items. While the methodological protocol itself is completely open, individual scales cited within the primer may require explicit authorization from original test developers, institutional copyright holders, or commercial publishers prior to diagnostic or clinical deployment.
12. References
- Boateng, G. O., Neilands, T. B., Frongillo, E. A., Melgar-Quiñonez, H. R., & Young, S. L. (2018). Best practices for developing and validating scales for health, social, and behavioral research: A primer. Frontiers in Public Health, 6, Article 149. https://doi.org/10.3389/fpubh.2018.00149
- Browne, M. W., & Cudeck, R. (1993). Alternative ways of assessing model fit. In K. A. Bollen & J. S. Long (Eds.), Testing Structural Equation Models (pp. 136–162). SAGE Publications.
- Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81–105. https://doi.org/10.1037/h0046016
- Cattell, R. B. (1966). The scree test for the number of factors. Multivariate Behavioral Research, 1(2), 245–276. https://doi.org/10.1207/s15327906mbr0102_10
- Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297–334. https://doi.org/10.1007/BF02310555
- DeVellis, R. F. (2012). Scale Development: Theory and Applications (3rd ed.). SAGE Publications.
- Horn, J. L. (1965). A rationale and test for the number of factors in factor analysis. Psychometrika, 30(2), 179–185. https://doi.org/10.1007/BF02289447
- Hu, L., & Bentler, P. M. (1999). Cutoff criteria for fit indexes in covariance structure analysis: Conventional criteria versus new alternatives. Structural Equation Modeling: A Multidisciplinary Journal, 6(1), 1–55. https://doi.org/10.1080/10705519909540118
- Lawshe, C. H. (1975). A quantitative approach to content validity. Personnel Psychology, 28(4), 563–575. https://doi.org/10.1111/j.1744-6570.1975.tb01393.x
- Lynn, M. R. (1986). Determination and quantification of content validity. Nursing Research, 35(6), 382–385. https://doi.org/10.1097/00006199-198611000-00017
- McDonald, R. P. (1999). Test Theory: A Unified Treatment. Lawrence Erlbaum Associates.
- Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741–749. https://doi.org/10.1037/0003-066X.50.9.741
- Nunnally, J. C., & Bernstein, I. H. (1994). Psychometric Theory (3rd ed.). McGraw-Hill.
- Streiner, D. L., Norman, G. R., & Cairney, J. (2015). Health Measurement Scales: A Practical Guide to Their Development and Use (5th ed.). Oxford University Press. https://doi.org/10.1093/med/9780199685219.001.0001
- Vandenberg, R. J., & Lance, C. E. (2000). A review and synthesis of the measurement invariance literature: Suggestions, practices, and recommendations for organizational research. Organizational Research Methods, 3(1), 4–70. https://doi.org/10.1177/109442810031002
13. Items of the Scale
The source document, “Best Practices for Developing and Validating Scales for Health, Social, and Behavioral Research: A Primer” (Boateng et al., 2018), is an overarching methodological guideline and psychometric blueprint rather than a single clinical questionnaire. The official, concrete measurement items derived from this methodology (such as specialized health insecurity questionnaires, behavioral scales, or clinical screening tools developed by the authors) are subject to individual author copyright, empirical field documentation, or institutional permissions, and the complete standardized item inventory is not reproduced in the open public domain.
To acquire the authorized, fully validated item inventories for specific empirical scales developed using this primer—such as the Household Water Insecurity Experiences (HWISE) Scale or related global health measures—researchers must contact the primary corresponding author (Dr. Godfred O. Boateng at Northwestern University / University of Texas at Arlington) or access the specific peer-reviewed study publications detailing those individual empirical instruments.
Illustrative Structural Implementation: The 9-Step Methodological Audit Checklist
The core structural product generated within the Boateng et al. (2018) primer is the systematic Nine-Step Methodological Scale Development and Validation Protocol. Below is the operational evaluation framework used by psychometricians to assess the methodological fidelity of newly developed scales:
Phase I: Item Development Audit
-
Step 1: Domain Identification & Item Generation
[ ] Clearly defined theoretical domain and construct boundaries.
[ ] Deductive literature review executed across relevant theoretical disciplines.
[ ] Inductive qualitative interviews / focus groups conducted with target demographic.
[ ] Initial pool generated with over-inclusive item volume (typically 2–3× final target).
[ ] Response scale options (Likert, semantic differential, dichotomous) explicitly established. -
Step 2: Content Validity Assessment
[ ] Expert review panel convened (minimum 5 to 10 independent subject-matter experts).
[ ] Quantitative Lawshe’s CVR or Lynn’s CVI calculated per item.
[ ] Items pruned or linguistically revised based on standardized CVR/CVI thresholds. -
Step 3: Cognitive Pre-Testing
[ ] Qualitative cognitive interviews (“think-aloud” protocol, verbal probing) administered.
[ ] Participant comprehension, recall burden, and response mapping formally evaluated.
[ ] Ambiguous, double-barreled, or culturally insensitive items modified or eliminated.
Phase II: Scale Development Audit
-
Step 4: Scale Administration & Survey Sampling
[ ] Adequate sample size recruited (target 5:1 to 10:1 subject-to-item ratio; N ≥ 200–300).
[ ] Sample recruited from representative target populations across study regions.
[ ] Missing data patterns examined; FIML or multiple imputation implemented if appropriate. -
Step 5: Item Reduction Procedures
[ ] Response distributions inspected for floor or ceiling effects (> 80% endorsement).
[ ] Corrected item-total correlations calculated; items with r < 0.30 removed.
[ ] Bivariate inter-item correlations inspected for redundancy (r > 0.70–0.80) or collinearity. -
Step 6: Latent Factor Extraction
[ ] Parallel Analysis and/or Velicer’s MAP test utilized to determine factor retention.
[ ] Exploratory factor analysis executed using robust extraction (ML or PAF).
[ ] Oblique factor rotation (e.g., Promax, Direct Quartimin) applied.
[ ] Factor loadings evaluated; items with loadings < 0.40 or cross-loadings removed.
Phase III: Scale Evaluation Audit
-
Step 7: Tests of Dimensionality & Verification
[ ] Independent hold-out validation sample utilized for structural verification.
[ ] Confirmatory Factor Analysis (CFA) or ESEM estimated.
[ ] Goodness-of-fit confirmed: CFI ≥ 0.95, TLI ≥ 0.95, RMSEA ≤ 0.06, SRMR ≤ 0.08.
[ ] Multigroup CFA conducted to substantiate configural, metric, and scalar invariance. -
Step 8: Scale Reliability Evaluation
[ ] Internal consistency estimated via McDonald’s omega coefficient (ω ≥ 0.70 to 0.80).
[ ] Ordinal alpha/theta computed if response data are skewed or categorical.
[ ] Longitudinal test-retest reliability estimated via two-way mixed ICC (target ≥ 0.75). -
Step 9: Construct & Criterion Validity Testing
[ ] Convergent validity verified against correlated theoretical measures (r ≥ 0.50).
[ ] Discriminant validity confirmed via Fornell-Larcker criteria or MTMM matrices.
[ ] Criterion-related validity documented via concurrent or prospective predictive modeling.