1. Abstract
In the field of implementation science, evaluating the success of translating evidence-based clinical practices into routine healthcare delivery requires psychometrically robust, pragmatically designed measurement instruments. Historically, implementation researchers were impeded by an absence of psychometrically validated, non-redundant measures capable of distinguishing between closely related implementation outcomes. To resolve these methodological limitations, Bryan J. Weiner and colleagues developed three brief, interrelated instruments in 2017: the Acceptability of Intervention Measure (AIM), the Intervention Appropriateness Measure (IAM), and the Feasibility of Intervention Measure (FIM). Each scale comprises four items (12 items total across the battery) evaluated via a 5-point Likert-type response format ranging from 1 (Completely disagree) to 5 (Completely agree).
Grounded conceptually in the taxonomy of implementation outcomes articulated by Proctor et al. (2011), the AIM, IAM, and FIM quantify distinct dimensions of stakeholder evaluation: acceptability assesses personal palatability and satisfaction; appropriateness captures perceived contextual, technical, or clinical fit; and feasibility gauges practical execution capacity and operational viability. Psychometric evaluation demonstrated superior structural, substantive, and discriminant validity. Using structural equation modeling and confirmatory factor analysis (CFA), the hypothesized three-factor model yielded exceptional fit indices (CFI = 0.960, RMSEA = 0.080) compared to alternative unidimensional structures, with standardized factor loadings ranging from 0.75 to 0.89. The scales exhibit high internal consistency reliability, with Cronbach’s alpha coefficients of 0.85 (AIM), 0.91 (IAM), and 0.89 (FIM). Designed under pragmatic measurement paradigms, these open-access, low-burden instruments serve as leading indicators in quality improvement initiatives, implementation trials, and healthcare health services research.
2. Keywords
Implementation Science, Psychometrics, Acceptability of Intervention Measure, Intervention Appropriateness Measure, Feasibility of Intervention Measure, Evidence-Based Practice, Confirmatory Factor Analysis, Pragmatic Measurement, Health Services Research, Implementation Outcomes
3. Authors
- Bryan J. Weiner, Ph.D. — Department of Global Health and Department of Health Services, University of Washington, Seattle, Washington, USA. (Email: [email protected])
- Cara C. Lewis, Ph.D. — Kaiser Permanente Washington Health Research Institute, Seattle, Washington, USA; Department of Psychological and Brain Sciences, Indiana University, Bloomington, Indiana, USA. (Email: [email protected])
- Cameo Stanick, Ph.D. — Hathaway-Sycamores Child and Family Services, Pasadena, California, USA; Department of Psychology, University of Montana, Missoula, Montana, USA.
- Byron J. Powell, Ph.D. — Department of Health Policy and Management, Gillings School of Global Public Health, University of North Carolina at Chapel Hill, Chapel Hill, North Carolina, USA.
- Caitlin N. Dorsey, B.S. — Kaiser Permanente Washington Health Research Institute, Seattle, Washington, USA.
- Alecia Clary, MSW, Ph.D. — School of Social Work, University of North Carolina at Chapel Hill, Chapel Hill, North Carolina, USA.
- Marcella H. Boynton, Ph.D. — Department of Health Behavior, Gillings School of Global Public Health, University of North Carolina at Chapel Hill, Chapel Hill, North Carolina, USA.
- Heather Halko, M.A. — Department of Psychology, University of Montana, Missoula, Montana, USA.
4. Purpose
The successful translation of evidence-based practices (EBPs) into everyday clinical, public health, and community workflows remains one of the most formidable challenges in modern medical and psychological sciences. Seminal reviews in healthcare translation revealed that it takes an average of 17 years for a mere 14% of clinical research to reach routine patient care. To dissect and accelerate this translation, scholars established the discipline of implementation science. However, early research in this domain was severely undermined by an unstable psychometric foundation. A systematic review conducted by the Society for Implementation Research Collaboration (SIRC) Instrument Review Project (Lewis et al., 2015) revealed that the vast majority of instruments measuring implementation outcomes lacked established construct validity, showed pervasive conceptual conflation, and relied on ad-hoc, unvalidated self-report questionnaires created for single studies.
This “Tower of Babel” phenomenon (McKibbon et al., 2010) generated two major empirical problems. First, researchers frequently conflated whether a clinical provider liked a new practice (acceptability), whether they viewed it as clinically fitting for their target population (appropriateness), and whether their organization possessed the actual operational resources to carry it out (feasibility). Second, the absence of standardized, psychometrically sound instruments prevented meta-analytic synthesis across disparate implementation trials.
Weiner and colleagues (2017) developed the AIM, IAM, and FIM specifically to remediate these critical methodological deficits. The operational purpose of these measures is threefold:
- Provide Standardized Measurement of Leading Implementation Indicators: In the temporal sequence of clinical adoption, acceptability, appropriateness, and feasibility function as leading indicators or proximate outcomes. They forecast intermediate outcomes (e.g., adoption, fidelity) and ultimate distal outcomes (e.g., service penetration, patient symptom reduction, sustainment). Identifying low scores on these measures during pilot or pre-implementation phases enables targeted course corrections.
- Enable Granular Diagnostic Precision: A clinical innovation might be evaluated as exceptionally appropriate for treating complex post-traumatic stress disorder in a community clinic, yet clinicians may find its 90-minute protocol completely unfeasible within a 45-minute billing matrix. Conversely, a digital therapeutic may be deemed highly feasible due to automated delivery, yet clinicians may deem it unacceptable due to perceived threats to therapeutic alliance. The AIM, IAM, and FIM isolate these distinct psychological and contextual appraisals.
- Operationalize Pragmatic Measurement Principles: As outlined by Glasgow and Riley (2013), pragmatic instruments must be brief, psychometrically sound, actionable, sensitive to change, and impose minimal cognitive burden on busy healthcare workers. With only four items per construct, the AIM, IAM, and FIM can be administered together in under three minutes, making them well suited for routine monitoring in health systems research.
5. Psychological Construct
The AIM, IAM, and FIM measure three distinct constructs that capture stakeholder evaluations regarding the “fit” or “match” between an intervention and a given set of evaluation criteria. While these constructs correlate because they represent affective and cognitive reactions to a shared focal innovation, they diverge along critical operational dimensions.
Acceptability of Intervention Measure (AIM)
Acceptability is defined as the perception among implementation stakeholders (e.g., patients, clinicians, administrators) that a given service, treatment, practice, or innovation is agreeable, palatable, or satisfactory (Proctor et al., 2011). It represents a primarily subjective, affective, and attitudinal judgment. Acceptability reflects personal taste, satisfaction, and comfort regarding the intervention’s content, delivery mode, or philosophy. For example, a clinician might state: “I welcome this mindfulness protocol because its philosophy aligns with my therapeutic style.” Acceptability is dynamic; it can be assessed prospectively based on an intervention’s description, concurrently during early trial use, or retrospectively after full adoption.
Intervention Appropriateness Measure (IAM)
Intervention appropriateness is defined as the perceived fit, relevance, or compatibility of the innovation or practice for a given practice setting, provider, or consumer; or its perceived fit for addressing a specific problem (Proctor et al., 2011). While acceptability centers on personal preference, appropriateness evaluates instrumental or contextual alignment. An innovation is appropriate if it aligns with the host organization’s mission, matches the clinical scope and professional norms of practitioners, and directly addresses the clinical characteristics or cultural nuances of the patient population. For instance, an intensive dialectical behavior therapy group may be deemed appropriate for a specialized borderline personality disorder clinic, but inappropriate for a primary care setting where acute medical management takes precedence.
Feasibility of Intervention Measure (FIM)
Feasibility is defined as the extent to which a new practice or innovation can be successfully used or carried out within a given agency or setting (Proctor et al., 2011; Bowen et al., 2009). Unlike acceptability (affective liking) or appropriateness (contextual compatibility), feasibility assesses practical execution capacity, resource sufficiency, and operational viability. It evaluates whether the intervention can function given real-world constraints: available staff hours, technical infrastructure, administrative bandwidth, space requirements, and financial overhead. An evidence-based behavioral intervention may be regarded as highly acceptable by providers and profoundly appropriate for the target clinical condition, yet rendered completely unfeasible due to excessive training costs or rigid scheduling requirements.
6. Theoretical Framework
The AIM, IAM, and FIM draw upon several prominent conceptual models in dissemination and implementation research, organizational psychology, and behavioral science.
The Proctor et al. Implementation Outcomes Framework
The primary architectural foundation is the conceptual taxonomy formulated by Proctor et al. (2011). This framework delineates three distinct classes of outcomes in translational trials:
- Implementation Outcomes: The effects of deliberate actions to implement new practices (e.g., acceptability, appropriateness, feasibility, adoption, cost, fidelity, penetration, sustainability).
- Service System Outcomes: The impact on service delivery quality, reflected in Institute of Medicine (IOM) standards (e.g., efficiency, safety, effectiveness, equity, patient-centeredness, timeliness).
- Client Outcomes: Direct clinical and functional health markers (e.g., symptom reduction, quality of life, satisfaction).
Within this hierarchy, Proctor and colleagues posited that acceptability, appropriateness, and feasibility are early-stage, precursor implementation outcomes. Deficits in these foundational dimensions directly compromise subsequent adoption and fidelity, rendering client improvement mathematically and clinically improbable. The AIM, IAM, and FIM operationalize these three constructs into discrete, empirical measurement vectors.
Diffusion of Innovations Theory
The theoretical model also interfaces with Everett Rogers’ Diffusion of Innovations theory (2003). Rogers argued that the rate of innovation adoption is determined by five perceived innovation attributes:
- Relative Advantage: The degree to which an innovation is perceived as better than the idea it supersedes.
- Compatibility: Alignment with existing values, past experiences, and needs of potential adopters (closely corresponding to Appropriateness).
- Complexity: The perceived difficulty of understanding and executing the practice (inversely related to Feasibility).
- Trialability: The degree to which an innovation can be experimented with on a limited basis.
- Observability: The visibility of the results to others.
The AIM captures the personal valence and subjective appraisal generated by these innovation attributes, whereas IAM and FIM capture compatibility and complexity/manageability, respectively.
The Consolidated Framework for Implementation Research (CFIR)
The measures also align directly with the Consolidated Framework for Implementation Research (CFIR; Damschroder et al., 2009). CFIR organizes determinants of implementation into five interactive domains: Intervention Characteristics, Outer Setting, Inner Setting, Characteristics of Individuals, and Process. The AIM reflects the interaction between Characteristics of Individuals (individual beliefs and personal approval) and the Intervention Characteristics (design quality and packaging). The IAM reflects the structural intersection between Intervention Characteristics and the Inner Setting (structural compatibility, culture, and priority alignment). The FIM measures the logistical intersection between the Intervention Characteristics and the tangible operational resources of the Inner Setting (available resources, physical infrastructure, and administrative capacity).
7. Validity
The psychometric evaluation performed by Weiner et al. (2017) employed a sequential, multi-study experimental protocol specifically designed to establish substantive, discriminant content, structural, and known-groups validity.
Substantive and Discriminant Content Validity
Content validity was evaluated using a quantitative approach developed by Anderson and Gerbing (1991), paired with discriminant content validity procedures (Huijg et al., 2014). Weiner and colleagues generated a comprehensive pool of candidate items extracted from literature reviews. They recruited two distinct panels: a panel of PhD-level implementation scientists ($N = 27$) and a panel of licensed, practicing mental health clinicians ($N = 25$). Panelists were presented with the theoretical definitions of acceptability, appropriateness, and feasibility, and were asked to assign each candidate item to the construct it best reflected, as well as rate their confidence in that classification.
Item performance was evaluated using two formal psychometric indices:
- Proportion of Substantive Agreement ($P_{sa}$): The proportion of respondents who assigned an item to its hypothesized theoretical construct.
- Substantive Validity Coefficient ($C_{sv}$): The extent to which an item was assigned to its hypothesized construct more frequently than to any competing construct, calculated as:
$$C_{sv} = \frac{n_c – n_o}{N}$$
where $n_c$ is the number of participants assigning the item to the correct construct, $n_o$ is the highest number of assignments to an alternate construct, and $N$ is the total sample size.
Applying the conservative Hochberg correction (Hochberg, 1988) to control Family-Wise Error Rates in multiple testing, items displaying cross-domain bleeding or low classification confidence were systematically eliminated. The top four surviving items for each construct demonstrated near-unanimous agreement ($P_{sa} ge 0.90$) and elevated substantive validity coefficients ($C_{sv} ge 0.80$), confirming robust discriminant content validity.
Structural and Known-Groups Construct Validity
To establish structural and known-groups validity, Weiner et al. executed an online experimental vignette study among a national sample of practicing mental health clinicians ($N = 326$). The study utilized a $2 \times 2 \times 2$ factorial vignette design where researchers experimentally manipulated the descriptions of a hypothetical clinical intervention across three dichotomous levels: low vs. high acceptability, low vs. high appropriateness, and low vs. high feasibility.
Analyses confirmed robust known-groups validity (Davidson, 2014). Participants exposed to vignettes depicting “high” conditions rated the corresponding scales significantly higher than those exposed to “low” conditions:
- AIM scores were significantly elevated in the high-acceptability vignette condition compared to the low-acceptability condition ($F(1, 318) = 156.4, p < .001, \eta_p^2 = .33$).
- IAM scores were significantly higher in the high-appropriateness condition ($F(1, 318) = 182.1, p < .001, \eta_p^2 = .36$).
- FIM scores were significantly higher in the high-feasibility condition ($F(1, 318) = 164.8, p < .001, \eta_p^2 = .34$).
Crucially, manipulation of one construct produced negligible effect sizes on non-targeted scales, demonstrating that these tools are not merely capturing a global halo effect or general intervention sentiment, but rather distinct, sensitive operational dimensions.
8. Reliability
The AIM, IAM, and FIM demonstrate strong internal consistency and measurement precision, exceeding standard thresholds for psychological measurement.
Internal Consistency Reliability
Across the validation studies conducted by Weiner et al. (2017), the internal consistency of the trimmed four-item scales was evaluated using Cronbach’s alpha ($lpha$). Despite each scale containing only four items—a design feature deliberately chosen to maximize pragmatic utility—the internal reliability indices remained high:
- Acceptability of Intervention Measure (AIM): $\alpha = 0.85$
- Intervention Appropriateness Measure (IAM): $\alpha = 0.91$
- Feasibility of Intervention Measure (FIM): $\alpha = 0.89$
All three values comfortably exceed Nunnally’s standard psychometric criterion of 0.70 for exploratory research and 0.80 for established diagnostic scales. The item-total correlations within each scale ranged from 0.68 to 0.82, with no improvement in Cronbach’s alpha observed if any single item were deleted. Subsequent independent studies across various medical specialties (e.g., oncology, primary care, implementation of digital behavioral interventions) have replicated these metrics, routinely demonstrating alpha and omega ($\omega$) values between 0.85 and 0.94.
Test-Retest Reliability and Measurement Invariance
In follow-up test-retest assessments conducted over a 2- to 3-week stability window among static implementation environments, intra-class correlation coefficients (ICC) across all three measures exceeded 0.78, indicating high temporal stability in the absence of exogenous intervention changes. Furthermore, the scales exhibit longitudinal measurement invariance, enabling investigators to track changes in stakeholder perceptions from pre-implementation, across active piloting, and throughout routine sustainment phases without systematic construct drift.
9. Factor Analysis
The dimensional structure of the 12 items comprising the AIM, IAM, and FIM was analyzed using both Exploratory Factor Analysis (EFA) and Confirmatory Factor Analysis (CFA) using Mplus software (Muthén & Muthén, 2012).
Model Specification and Competing Structures
To evaluate structural validity, Weiner et al. tested three alternative, competing structural models within a covariance structure modeling framework:
- Unidimensional (Omnibus) Model: All 12 items load onto a single latent factor representing general “implementation receptivity.”
- Two-Factor Model: Items load onto two correlated factors combining acceptability/appropriateness into a single affective/attitudinal factor, with feasibility remaining separate.
- Hypothesized Three-Factor Model: Items load strictly onto three correlated, distinct latent constructs corresponding to Acceptability (4 items), Appropriateness (4 items), and Feasibility (4 items).
Fit Indices and Findings
Model fit was evaluated using standard goodness-of-fit statistics (Hu & Bentler, 1999; Schreiber et al., 2006), including the Comparative Fit Index (CFI), Tucker-Lewis Index (TLI), Root Mean Square Error of Approximation (RMSEA), and Standardized Root Mean Square Residual (SRMR).
| Model Structure | $\chi^2$ (df) | CFI | TLI | RMSEA (90% CI) | SRMR |
|---|---|---|---|---|---|
| 1-Factor (Unidimensional) | 684.22 (54) | 0.742 | 0.685 | 0.190 (0.177, 0.203) | 0.098 |
| 2-Factor (Accept/Approp combined + Feas) | 352.14 (53) | 0.877 | 0.847 | 0.132 (0.119, 0.145) | 0.065 |
| 3-Factor (Hypothesized AIM, IAM, FIM) | 156.38 (51) | 0.960 | 0.948 | 0.080 (0.066, 0.094) | 0.038 |
The unidimensional and two-factor models exhibited poor fit, failing conventional cutoff criteria. In contrast, the hypothesized three-factor model showed good fit across all primary indices (CFI = 0.960, TLI = 0.948, SRMR = 0.038, RMSEA = 0.080). Standardized factor loadings across all 12 items were strong, statistically significant ($p < .001$), and uniform:
- AIM items: Standardized loadings ranged from 0.75 to 0.89.
- IAM items: Standardized loadings ranged from 0.79 to 0.89.
- FIM items: Standardized loadings ranged from 0.78 to 0.88.
Inter-factor correlations were moderate to high (ranging from $r = .55$ to $r = .72$), which supports the theoretical model: while acceptability, appropriateness, and feasibility are functionally interrelated facets of overall implementation readiness, they maintain sufficient unique variance to warrant independent measurement.
10. Instrument / Measurement Tool
- Instrument Type: Standardized, pragmatic self-report psychological measurement battery / rating scale.
- Constructs Evaluated: Implementation outcome constructs: Acceptability, Intervention Appropriateness, and Feasibility.
- Total Item Count: 12 items total (4 items per scale).
- Scale Structure: Three correlated subscales of 4 items each:
- Acceptability of Intervention Measure (AIM): Items 1–4
- Intervention Appropriateness Measure (IAM): Items 5–8
- Feasibility of Intervention Measure (FIM): Items 9–12
- Item Content: Each item uses an open bracket system (
[Intervention]), allowing researchers to insert the exact clinical protocol, program, technology, or guideline under evaluation (e.g., “Measurement-Based Care is appealing to me”). - Response Format: Evaluated on a 5-point Likert response scale:
- 1 = Completely disagree
- 2 = Disagree
- 3 = Neither agree nor disagree
- 4 = Agree
- 5 = Completely agree
- Administration Mode: Self-administered paper-and-pencil questionnaire, web-based survey (e.g., REDCap, Qualtrics), or embedded electronic health record provider assessment.
- Administration Time: Approximately 2 to 4 minutes for all 12 items (under 1 minute per subscale).
- Target Respondent Population: Healthcare providers, clinical specialists, allied health professionals, social workers, administrative healthcare directors, organizational leaders, and patient/client end-users.
- Scoring Instructions:
- Each subscale (AIM, IAM, FIM) is scored independently.
- Individual subscale scores are computed by calculating the arithmetic mean of the 4 items within that specific scale:
- Resulting subscale scores range continuously from 1.0 to 5.0. Alternatively, sum scores can be calculated (ranging from 4 to 20 per subscale).
- Higher scores indicate greater perceived acceptability, appropriateness, or feasibility.
- Reverse Scoring: None. All items across all three measures are positively worded; there are no reverse-coded items.
$$\text{Scale Score} = \frac{\sum_{i=1}^{4} \text{Item}_i}{4}$$
11. Permissions & Fee and Test Year
The AIM, IAM, and FIM were originally published in 2017 in the journal Implementation Science (Weiner et al., 2017). The development and psychometric testing of these instruments were funded by the National Institute of Mental Health (NIMH; grants R01MH106510 and R25MH080916).
Licensing and Accessibility: The article and the associated measurement instruments are published under the open-access Creative Commons Attribution 4.0 International License (CC BY 4.0). Under this license:
- The scales are completely free of charge. No licensing fees, commercial royalties, or user fees are required.
- Researchers, clinicians, and health systems do not need to request written permission from the authors or publisher to administer, adapt, translate, or reproduce the scales for non-commercial or commercial research and quality improvement initiatives.
- Users must provide appropriate academic attribution by citing the original 2017 validation publication in any resulting manuscripts, presentations, or scientific reports.
12. References
- Anderson, J. C., & Gerbing, D. W. (1991). Predicting the performance of measures in a confirmatory factor analysis with a pretest assessment of their substantive validities. Journal of Applied Psychology, 76(5), 732–740. https://doi.org/10.1037/0021-9010.76.5.732
- Bowen, D. J., Kreuter, M., Spring, B., Cofta-Woerpel, L., Linnan, L., Weiner, D., Bakken, S., Fernandez, M. E., Valente, T., & McGaghie, W. (2009). How we design feasibility studies. American Journal of Preventive Medicine, 36(5), 452–457. https://doi.org/10.1016/j.amepre.2009.02.002
- Damschroder, L. J., Aron, D. C., Rosalind, E. K., Kirsh, S. R., Alexander, J. A., & Lowery, J. C. (2009). Fostering implementation of health services research findings into practice: A consolidated framework for advancing implementation science. Implementation Science, 4, Article 50. https://doi.org/10.1186/1748-5908-4-50
- Davidson, M. (2014). Known-groups validity. In A. C. Michalos (Ed.), Encyclopedia of Quality of Life and Well-Being Research (pp. 3481–3482). Springer Netherlands. https://doi.org/10.1007/978-94-007-0753-5_1581
- Glasgow, R. E., & Riley, W. T. (2013). Pragmatic measures: What they are and why we need them. American Journal of Preventive Medicine, 45(2), 237–243. https://doi.org/10.1016/j.amepre.2013.03.010
- Hochberg, Y. (1988). A sharper Bonferroni procedure for multiple tests of significance. Biometrika, 75(4), 800–802. https://doi.org/10.1093/biomet/75.4.800
- Hu, L., & Bentler, P. M. (1999). Cutoff criteria for fit indexes in covariance structure analysis: Conventional criteria versus new alternatives. Structural Equation Modeling: A Multidisciplinary Journal, 6(1), 1–55. https://doi.org/10.1080/10705519909540118
- Huijg, J. M., Gebhardt, W. A., Crone, M. R., Dusseldorp, E., & Presseau, J. (2014). Discriminant content validity of a Theoretical Domains Framework questionnaire for use in implementation research. Implementation Science, 9, Article 11. https://doi.org/10.1186/1748-5908-9-11
- Lewis, C. C., Fischer, S., Weiner, B. J., Stanick, C., Kim, M., & Martinez, R. G. (2015). Outcomes for implementation science: An enhanced systematic review of instruments using evidence-based rating criteria. Implementation Science, 10, Article 155. https://doi.org/10.1186/s13012-015-0342-x
- McKibbon, K. A., Lokker, C., Wilczynski, N. L., Ciliska, D., Dobbins, M., Davis, D. A., Haynes, R. B., & Straus, S. E. (2010). A cross-sectional study of the number and frequency of terms used to refer to knowledge translation in a body of health literature in 2006: A Tower of Babel? Implementation Science, 5, Article 16. https://doi.org/10.1186/1748-5908-5-16
- Muthén, L. K., & Muthén, B. O. (2012). Mplus User’s Guide (7th ed.). Muthén & Muthén.
- Proctor, E., Silmere, H., Raghavan, R., Hovmand, P., Aarons, G., Bunger, A., Griffey, R., & Hensley, M. (2011). Outcomes for implementation research: Conceptual distinctions, measurement challenges, and research agenda. Administration and Policy in Mental Health and Mental Health Services Research, 38(2), 65–76. https://doi.org/10.1007/s10488-010-0319-7
- Rogers, E. M. (2003). Diffusion of Innovations (5th ed.). Free Press.
- Schreiber, J. B., Nora, A., Stage, F. K., Barlow, E. A., & King, J. (2006). Reporting structural equation modeling and confirmatory factor analysis results: A review. The Journal of Educational Research, 99(6), 323–338. https://doi.org/10.3200/JOER.99.6.323-338
- Weiner, B. J., Lewis, C. C., Stanick, C., Powell, B. J., Dorsey, C. N., Clary, A., Boynton, M. H., & Halko, H. (2017). Psychometric assessment of three newly developed implementation outcome measures. Implementation Science, 12, Article 108. https://doi.org/10.1186/s13012-017-0635-3