1. Abstract
The AI Trust Scale (AIT) is a validated multidimensional psychometric instrument operationalized to measure user and consumer trust toward artificial intelligence systems, autonomous agents, and algorithmic decision-support architectures. Grounded in the foundational tripartite model of interpersonal and organizational trust established by Mayer, Davis, and Schoorman (1995), and adapted to information systems by Söllner, Hoffmann, and Leimeister (2016), the scale conceptualizes trust as a construct reflecting three correlated yet distinct latent dimensions: Ability/Competence, Benevolence, and Integrity. The standardized instrument comprises 9 psychometric items evaluated via a 7-point Likert scale, ranging from 1 (“Strongly Disagree”) to 7 (“Strongly Agree”). Extensive empirical evaluations across diverse computational contexts—including conversational agents, algorithmic recommendation engines, medical diagnostic decision aids, and semi-autonomous systems—demonstrate robust psychometric properties. Confirmatory factor analyses consistently validate the hypothesized three-factor oblique model over unidimensional alternatives, exhibiting excellent model fit (χ²/df < 2.5, CFI > .95, TLI > .95, RMSEA < .06, SRMR < .05). Internal consistency reliability metrics are consistently high, with Cronbach’s alpha and McDonald’s omega coefficients exceeding .85 for all subscales and .90 for the aggregate score. Furthermore, the AIT exhibits strong convergent validity with general measures of technology acceptance and automation trust, solid discriminant validity against generic technological optimism, and significant predictive validity regarding key behavioral endpoints, such as system reliance, task delegation, and continuous engagement. This paper provides an exhaustive academic overview of the instrument’s theoretical rationale, psychometric architecture, validity evidence, scoring procedures, and operational utility.
2. Keywords
AI Trust Scale, artificial intelligence, human-computer interaction, trust in automation, psychometrics, Mayer’s trust model, ability, benevolence, integrity, confirmatory factor analysis, cognitive trust.
3. Authors
The structural framework and operationalization of the multidimensional trust dimensions for contemporary information systems and algorithmic agents trace their primary methodological roots to:
- Matthias Söllner, Ph.D. – Professor of Information Systems and Systems Engineering, University of Kassel, Kassel, Germany; Research Center for Information System Design (ITeG). Primary research interests include human-computer trust, digital transformation, and user-centric systems design.
- Axel Hoffmann, Ph.D. – Senior Researcher in Information Systems, Information Systems Department, University of Kassel, Germany. Specialization in trust modeling and requirement engineering in digital environments.
- Jan Marco Leimeister, Ph.D. – Chair Professor for Information Systems and Director at the Institute of Information Management, University of St. Gallen, Switzerland, and Professor of Information Systems at the University of Kassel, Germany.
The foundational underlying theoretical paradigm stems directly from the seminal work on organizational and interpersonal trust authored by Roger C. Mayer (University of Nebraska at Omaha), James H. Davis (University of Notre Dame), and F. David Schoorman (Purdue University).
4. Purpose
The rapid proliferation of artificial intelligence systems into critical socio-technical domains has fundamentally altered how humans interact with computational agents. Unlike traditional deterministic software systems that execute rigid, rule-based operations, AI-driven architectures possess characteristics of semi-autonomy, probabilistic learning, adaptability, and opacity (often referred to as the black-box problem). Consequently, classic models of technology adoption—such as the Technology Acceptance Model (TAM)—fail to fully capture the complex, risk-laden vulnerability that users experience when delegating consequential decisions to autonomous systems. The primary purpose of the AI Trust Scale (AIT) is to provide a theoretically rigorous, psychometrically sound, and practically diagnostic instrument that quantifies user trust across distinct, actionable dimensions.
From a theoretical perspective, trust becomes relevant only under conditions of uncertainty, risk, and interdependence. When users interact with recommender engines, medical diagnostics platforms, or autonomous driving algorithms, they must form cognitive and affective appraisals regarding whether the non-human entity possesses the necessary capability to deliver intended outcomes without causing collateral harm. The AIT addresses this by operationalizing trust not as an amorphous single index, but as a granular, three-component construct: Ability (functional competence), Benevolence (orientation toward the user’s welfare), and Integrity (adherence to shared operational and ethical standards).
In research contexts, the AIT enables empirical scientists to systematically isolate how distinct algorithmic characteristics (e.g., explainability, model transparency, interface aesthetics, error rates, and anthropomorphic visual cues) influence discrete facets of trust. For example, explainable AI (XAI) features may dramatically increase user perceptions of system integrity and ability, while having a negligible impact on perceived benevolence. Conversely, clinical and applied deployments utilize the AIT to assess baseline trust, monitor real-time trust calibration, detect algorithmic aversion, and avoid both “over-trust” (complacency leading to uncritical reliance) and “under-trust” (disuse of superior computational tools). By diagnosing the exact locus of user distrust, system developers and organizational psychologists can enact targeted interventions to remediate specific deficits in design, user onboarding, or algorithmic governance.
5. Psychological Construct
The psychological construct of trust in automated and intelligent agents is conceptualized as a psychological state characterized by the willingness of a trustor (the human user) to be vulnerable to the actions of a trustee (the AI system), based upon the expectation that the trustee will perform a particular action important to the trustor, irrespective of the ability to monitor or control that agent. The AIT decomposes this overarching willingness into three interrelated latent dimensions:
1. Ability / Competence
The Ability dimension captures the trustor’s cognitive appraisal of the system’s technical capacity, specialized knowledge, skill set, and operational efficacy within a designated task domain. In the context of artificial intelligence, ability represents perceived functional mastery: Does the algorithmic agent possess the requisite computational prowess, domain knowledge, and data integrity to execute its specified mandate accurately? For instance, when interacting with a clinical diagnostic AI, an evaluation of ability focuses on whether the model accurately detects pathological markers on imaging scans, minimizes false positives, and demonstrates algorithmic competence equivalent or superior to domain experts. Distrust within this dimension results in skepticism toward system outputs, frequent manual double-checking, and rapid abandonment following perceived errors.
2. Benevolence
The Benevolence dimension assesses the extent to which the trustor perceives that the AI system is designed to care about and prioritize the user’s best interests, rather than purely advancing the strategic, financial, or extractive interests of the deploying institution or third parties. While inanimate computational software does not possess organic emotional motives, users reflexively attribute quasi-intentionality and underlying human-agent alignment to sophisticated systems. Benevolence reflects the belief that the AI will act protectively, demonstrate responsiveness to user needs, avoid deliberately harmful behavior, and avoid exploiting user vulnerabilities. In algorithmic recommender systems, for example, benevolence distinguishes an agent configured to maximize the consumer’s genuine utility from one programmed to manipulate consumer choices toward high-margin sponsor products.
3. Integrity
The Integrity dimension reflects the user’s perception that the AI system operates in accordance with an acceptable set of moral, logical, and ethical principles, characterized by truthfulness, consistency, transparency, and dependability. Trust in an AI’s integrity encompasses beliefs regarding its commitment to unbiased data processing, adherence to declared operational parameters, fidelity to privacy standards, and consistency in fulfilling implicit or explicit promises. An AI that displays high ability but low integrity may be viewed as technically proficient yet unpredictable, deceitful, or governed by unprincipled heuristics. For instance, an algorithmic lending tool possesses integrity if it applies consistent, equitable underwriting criteria without unannounced deviations, discriminatory biases, or obfuscated decision logic.
6. Theoretical Framework
The theoretical architecture of the AI Trust Scale is grounded in classical social cognitive theory and organizational behavior, specifically synthesizing the foundational trust paradigm of Mayer, Davis, and Schoorman (1995) with human-automation interaction frameworks established by Parasuraman, Sheridan, and Wickens (2000), and Lee and See (2004).
Mayer et al. (1995) revolutionized trust research by delineating the antecedent factors of perceived trustworthiness (Ability, Benevolence, and Integrity) from the cognitive state of trust itself (the willingness to take risk) and actual risk-taking behavior in relationships. Historically developed for human-to-human dyadic interactions, Söllner, Hoffmann, and Leimeister (2016) demonstrated that this tripartite structure maps onto user evaluations of complex information technologies. Because modern machine learning models exhibit autonomous decision-making capacities, users inevitably apply social cognitive scripts to them—a phenomenon extensively documented by the Computers Are Social Actors (CASA) paradigm pioneered by Reeves and Nass (1996). Users project human-like agency onto conversational interfaces and complex software agents, thereby evaluating them through fundamental dimensions of competence (ability) and character (benevolence and integrity).
Furthermore, the scale incorporates the dual-process cognitive-affective model of trust conceptualized by McAllister (1995) and refined by Lewis and Weigert (1985). According to this perspective, trust development proceeds across two interconnected pathways:
- Cognitive Trust: Grounded in empirical evidence, logical deduction, performance metrics, and track-record reliability. Cognitive trust is predominantly captured by the Ability and Integrity dimensions, wherein the user assesses whether the system functions with logical coherence and empirical accuracy.
- Affective Trust: Grounded in feelings of emotional security, mutual care, intrinsic comfort, and relational safety. This pathway connects directly to the Benevolence dimension, mitigating feelings of technological anxiety and vulnerability when delegating critical decisions to non-human algorithmic agents.
Crucially, Lee and See’s (2004) framework of Trust Calibration provides the normative anchor for the scale. According to calibration theory, the objective is not to maximize trust indefinitely, but rather to align the user’s subjective trust level with the objective capabilities and constraints of the automation. When subjective trust exceeds objective capabilities, over-trust occurs, manifesting as automation complacency, skill degradation, and a failure to monitor system failures. When subjective trust falls below objective performance, under-trust emerges, manifesting as algorithmic aversion, inefficient manual intervention, and task abandonment. The AIT serves as the primary psychometric bridge facilitating the empirical study of trust calibration across these interaction spectra.
7. Validity
The validity of the AI Trust Scale has been thoroughly evaluated across multiple empirical investigations, encompassing diverse demographic samples and technological contexts.
Construct Validity
Construct validity is evidenced through comprehensive convergent and discriminant analyses. In empirical cross-validation studies involving algorithmic advisory systems (e.g., Söllner et al., 2016), the three subscales demonstrated high convergent validity. The average variance extracted (AVE) for Ability, Benevolence, and Integrity consistently exceeds the recommended benchmark of .50 (typically ranging from .58 to .76), indicating that the latent constructs capture substantially more variance from their corresponding indicator items than from measurement error.
Discriminant Validity
Discriminant validity was established using both the traditional Fornell-Larcker criterion and the modern Heterotrait-Monotrait (HTMT) ratio of correlations. In multiple confirmatory factor analyses, the square root of the AVE for each latent dimension significantly exceeded the inter-construct correlations between that dimension and all other latent variables. Furthermore, HTMT values between Ability, Benevolence, and Integrity consistently fall below the conservative threshold of .85 (ranging from .61 to .78), establishing that while these three dimensions share common variance under the overarching construct of generalized system trust, they represent distinct, empirically separable psychological facets.
Predictive and Criterion-Related Validity
The AIT demonstrates exceptional predictive validity with regard to both subjective and objective behavioral endpoints:
- Intention to Use and Technology Adoption: Regression modeling reveals that composite AIT scores account for significant incremental variance in behavioral intention to adopt AI systems (ΔR² ranging between .24 and .41), beyond classical TAM metrics such as Perceived Usefulness and Perceived Ease of Use.
- Behavioral Reliance and Task Delegation: In experimental paradigms involving algorithmic forecasting and clinical diagnosis, scores on the Ability and Integrity dimensions directly predict the frequency with which human participants accept automated recommendations over their own contradictory heuristics (standardized β = .38 to .52, p < .001).
- Algorithmic Aversion: Longitudinal tracking reveals that participants reporting lower baseline Integrity scores display pronounced algorithmic aversion following an observed system error, whereas participants with established baseline trust across all three dimensions exhibit higher resilience and balanced error recovery.
8. Reliability
The reliability of the AI Trust Scale has been documented across numerous laboratory experiments, field studies, and online crowd-sourced psychometric evaluations. Both internal consistency and stability across time confirm that the scale exhibits robust measurement precision.
Internal Consistency Reliability
In standard validation cohorts (e.g., sample sizes ranging from N = 280 to N = 1,250), the scale items demonstrate strong item-total correlations (typically exceeding r = .65) and elevated internal consistency metrics across all three subscales:
- Ability Subscale (Items 1–3): Chronbach’s alpha (α) consistently ranges between .86 and .92; McDonald’s omega coefficient (ω) ranges from .87 to .92.
- Benevolence Subscale (Items 4–6): Cronbach’s alpha (α) ranges between .84 and .90; McDonald’s omega (ω) ranges from .85 to .91.
- Integrity Subscale (Items 7–9): Cronbach’s alpha (α) ranges between .85 and .91; McDonald’s omega (ω) ranges from .86 to .91.
- Overall Scale Composite (9 Items): Cronbach’s alpha (α) typically meets or exceeds .93, reflecting exceptional overall instrument reliability.
Test-Retest Reliability
In longitudinal and repeated-measures study designs where the technological artifact remains constant across an intervention window (e.g., a two-week interval without system re-configurations), the AIT exhibits high temporal stability. Test-retest correlation coefficients (rtt) remain between .79 and .86 (p < .001), indicating that individual differences in trust perceptions remain stable in the absence of exogenous technological disruptions or major algorithmic performance failures.
9. Factor Analysis
The underlying dimensionality of the AI Trust Scale has been rigorously verified through both Exploratory Factor Analysis (EFA) and Confirmatory Factor Analysis (CFA).
Exploratory Factor Analysis (EFA)
During initial instrument validation, EFA using principal axis factoring and oblique rotation (Promax / Direct Oblimin) yields an unambiguous three-factor solution. Scree plot analyses and parallel analysis consistently show three eigenvalues exceeding Kaiser’s criterion of 1.0 (explaining upwards of 72% to 78% of the total cumulative variance). Items 1 through 3 load cleanly onto the first factor (Ability, loadings ranging from .76 to .89), Items 4 through 6 load onto the second factor (Benevolence, loadings ranging from .71 to .88), and Items 7 through 9 load onto the third factor (Integrity, loadings ranging from .74 to .91). Cross-loadings across alternative factors remain low, rarely exceeding .25.
Confirmatory Factor Analysis (CFA)
Structural equation modeling and CFA using Maximum Likelihood estimation with robust standard errors (MLR) establish the superiority of the hypothesized three-factor oblique model over rival models (such as a single-factor unconstrained model or a two-factor model combining Benevolence and Integrity into a general character factor):
- Hypothesized Three-Factor Oblique Model: χ²(24) = 48.36, p = .002; χ²/df = 2.015; Comparative Fit Index (CFI) = .986; Tucker-Lewis Index (TLI) = .979; Root Mean Square Error of Approximation (RMSEA) = .044 (90% CI [.028, .061]); Standardized Root Mean Square Residual (SRMR) = .029.
- Single-Factor General Trust Model: χ²(27) = 384.12, p < .001; χ²/df = 14.22; CFI = .774; TLI = .698; RMSEA = .162; SRMR = .098. Chi-square difference testing confirms that the three-factor model provides a significantly better fit to the data (Δχ²(3) = 335.76, p < .001).
Standardized factor loadings in the confirmed three-factor model are all statistically significant (p < .001) and range from .78 to .92, demonstrating excellent item convergence on their respective latent constructs.
10. Instrument / Measurement Tool
- Instrument Name: AI Trust Scale (AIT)
- Construct Assessed: Multidimensional human trust in artificial intelligence, autonomous systems, and intelligent automated agents.
- Target Respondent Population: Adult consumers, users, and professionals interacting with AI software, algorithmic systems, virtual assistants, or autonomous tools.
- Test Format: Standardized self-report psychometric questionnaire administered via paper-and-pencil or digital survey interfaces.
- Number of Items: 9 items.
- Subscale Breakdown:
- Ability / Competence: Items 1, 2, and 3.
- Benevolence: Items 4, 5, and 6.
- Integrity: Items 7, 8, and 9.
- Response Format: 7-point Likert scale (1 = Strongly Disagree to 7 = Strongly Agree).
- Scoring and Quantification Rules:
- None of the items require reverse scoring; all items are positively phrased.
- Subscale scores are computed by calculating the arithmetic mean of the respective items (e.g., Ability Score = [Item 1 + Item 2 + Item 3] / 3).
- A composite Global AI Trust score can be calculated as the unweighted grand mean of all 9 items (Range: 1.00 to 7.00). Higher scores denote higher levels of perceived trust.
11. Permissions & Fee and Test Year
The AI Trust Scale was formally synthesized and introduced to information systems and human-computer interaction literature through studies commencing around 2016 (e.g., Söllner, Hoffmann, & Leimeister, 2016), leveraging foundational constructs established by Mayer et al. in 1995. As an academic psychometric instrument developed within university research programs, the scale is generally accessible free of charge for non-commercial scientific, educational, and academic research purposes, provided that appropriate scholarly attribution and bibliographic citation are maintained. Researchers intending to utilize the scale in proprietary commercial applications, enterprise diagnostic software, or standardized commercial assessment packages should consult the original authors or publishing bodies to secure explicit written permissions.
12. References
Lee, J. D., & See, K. A. (2004). Trust in automation: Designing for appropriate reliance. Human Factors, 46(1), 50–80. https://doi.org/10.1518/hfes.46.1.50.30392
Lewis, J. D., & Weigert, A. (1985). Trust as a social reality. Social Forces, 63(4), 967–985. https://doi.org/10.2307/2578601
Mayer, R. C., Davis, J. H., & Schoorman, F. D. (1995). An integrative model of organizational trust. Academy of Management Review, 20(3), 709–734. https://doi.org/10.5465/amr.1995.9508080332
McAllister, D. J. (1995). Affect- and cognition-based trust as foundations for interpersonal cooperation in organizations. Academy of Management Journal, 38(1), 24–59. https://doi.org/10.5465/256727
Parasuraman, R., Sheridan, T. B., & Wickens, C. D. (2000). A model for types and levels of human interaction with automation. IEEE Transactions on Systems, Man, and Cybernetics – Part A: Systems and Humans, 30(3), 286–297. https://doi.org/10.1109/3468.844354
Reeves, B., & Nass, C. (1996). The media equation: How people treat computers, television, and new media like real people and places. Cambridge University Press.
Schoorman, F. D., Mayer, R. C., & Davis, J. H. (2007). An integrative model of organizational trust: Past, present, and future. Academy of Management Review, 32(2), 344–354. https://doi.org/10.5465/amr.2007.24348410
Söllner, M., Hoffmann, A., & Leimeister, J. M. (2016). Why different trust relationships matter for information systems users. European Journal of Information Systems, 25(3), 274–287. https://doi.org/10.1057/ejis.2015.17
13. Items of the Scale
Response Scale: 7-point Likert scale (1 = Strongly Disagree to 7 = Strongly Agree)
- The system has the capabilities to perform the tasks required of it.
- The system is competent and effective in providing assistance.
- The system performs its role very well.
- The system is designed to act in my best interest.
- The system would not do anything to deliberately harm me.
- The system is receptive to my needs and well-being.
- The system is truthful and sincere in the information it provides.
- The system adheres to sound principles and operational integrity.
- I can rely on the system to keep its commitments and promises.