Abstract
The ChatGPT Usage Scale is a specialized, empirically validated psychometric instrument developed to assess and quantify how postgraduate students integrate conversational artificial intelligence—specifically ChatGPT—into their academic and research workflows. Formulated against the backdrop of rapid advancements in generative artificial intelligence and large language models (LLMs), this 15-item self-report questionnaire transitions the assessment of academic AI adoption from anecdotal observation to rigorous quantitative measurement. Grounded conceptually in the Technology Acceptance Model (TAM) and Cognitive Load Theory (CLT), the scale measures an overarching multidimensional construct operationalized across three correlated subscales: Academic Writing Aid (6 items), Academic Task Support (5 items), and Reliance and Trust (4 items). Evaluated using a 5-point Likert response scale ranging from 1 (Strongly Disagree) to 5 (Strongly Agree), the instrument was standardized on a sample of 443 Egyptian postgraduate students across diploma, master’s, and doctoral programs.
Psychometric evaluation demonstrated robust structural integrity and internal consistency. Following an exploratory factor analysis (EFA) that reduced an initial 39-item pool to 15 items accounting for approximately 49% of the total variance, confirmatory factor analysis (CFA) verified acceptable structural fit indices: Comparative Fit Index (CFI) = 0.917, Tucker-Lewis Index (TLI) = 0.900, and Root Mean Square Error of Approximation (RMSEA) = 0.060. The scale demonstrates excellent reliability with a Cronbach’s alpha (α) of 0.848, McDonald’s omega (ω) of 0.849, and composite reliability (CR) of 0.855. Convergent validity was substantiated by an Average Variance Extracted (AVE) of 0.664, exceeding the conventional 0.50 threshold. By capturing mechanistic drafting behaviors, ideation support, and cognitive-psychological reliance, the ChatGPT Usage Scale serves as a crucial diagnostic and empirical tool for higher education institutions, educational psychologists, and academic policymakers navigating the balance between AI-assisted learning and academic integrity.
Keywords
ChatGPT Usage Scale, Generative Artificial Intelligence, Psychometrics, Higher Education, Postgraduate Students, Technology Acceptance Model, Cognitive Load Theory, Academic Integrity, Reliance and Trust, Factor Analysis
Authors
The ChatGPT Usage Scale was developed and validated by an academic research team from the Faculty of Education at Al-Azhar University, Egypt:
- Mohamed Nemt-Allah (Corresponding Author) — Department of Educational Psychology, Faculty of Education, Al-Azhar University, Dakahlia, Egypt. Email: [email protected]
- Waleed Khalifa — Faculty of Education, Al-Azhar University, Dakahlia, Egypt.
- Mahmoud Badawy — Faculty of Education, Al-Azhar University, Dakahlia, Egypt.
- Yasser Elbably — Faculty of Education, Al-Azhar University, Dakahlia, Egypt.
- Ashraf Ibrahim — Faculty of Education, Al-Azhar University, Dakahlia, Egypt.
Purpose
The rapid proliferation of generative artificial intelligence, particularly conversational agents driven by autoregressive neural networks such as OpenAI’s ChatGPT, has precipitated a paradigm shift across international higher education. While digital tools have long assisted students with reference management and spell-checking, generative AI represents a qualitative leap: it can synthesize dense literature, execute complex conceptual reframing, structure empirical research plans, generate computer code, and produce coherent academic discourse across virtually any discipline. Prior to the development of the ChatGPT Usage Scale, the scholarly community lacked an instrument calibrated to assess how advanced students—specifically postgraduate scholars—interact with these tools. Postgraduate education entails demanding cognitive outputs, including dissertation drafting, systematic peer-reviewed publication, advanced theoretical synthesis, and independent critical thinking. Consequently, generic measures of technology adoption failed to capture the multifaceted cognitive and ethical dynamics at play.
The ChatGPT Usage Scale was specifically engineered to fill this methodological void. The primary objective is to deliver a standardized, psychometrically grounded instrument that measures not merely binary adoption rates (e.g., whether a student uses AI or the frequency of use), but the functional and psychological nature of student engagement. Researchers, educators, and institutional leaders require empirical insight into whether students deploy ChatGPT as an intellectual scaffold to augment cognitive bandwidth, or whether they exhibit uncritical epistemic reliance that could undermine scholastic development and academic integrity.
In educational psychology and institutional research, the scale serves several essential clinical and investigative purposes:
- Diagnosing Cognitive Offloading and AI Dependence: It allows researchers to quantify cognitive offloading behaviors, identifying where supportive automation transitions into maladaptive cognitive dependence that impairs self-regulated learning and critical problem-solving skills.
- Correlational and Longitudinal Investigation: The tool enables researchers to evaluate how distinct modalities of AI usage correlate with academic anxiety, burnout, epistemic curiosity, academic self-efficacy, and objective academic performance metrics (such as thesis quality or GPA).
- Informing Evidence-Based Institutional AI Policies: Higher education administrators can utilize the instrument to map longitudinal baseline trends across academic departments, thereby formulating nuanced pedagogical guidelines that differentiate between legitimate academic scaffolding and academic misconduct.
- Evaluating Educational Interventions: Curriculum designers can employ the scale in pre-test/post-test experimental designs to evaluate whether explicit training in AI literacy, prompt engineering, and critical verification reduces uncritical reliance while preserving beneficial research productivity.
Psychological Construct
The overarching construct captured by the ChatGPT Usage Scale is Academic AI Engagement—specifically operationalized as the behavioral, cognitive, and affective patterns exhibited by postgraduate scholars when interacting with generative AI interfaces during the research and study lifecycle. Rather than treating technology use as a monolithic, one-dimensional frequency metric, the construct is conceptualized as a multidimensional, hierarchical domain. This captures functional operational utility alongside metacognitive trust and epistemic delegation. Grounded in psychometric evaluation, this overarching construct is composed of three interrelated sub-dimensions:
1. Academic Writing Aid
The Academic Writing Aid dimension captures the direct operationalization of ChatGPT as an automated drafting, stylistic revision, and rhetorical synthesis assistant. In postgraduate education, academic writing represents one of the most cognitively demanding tasks, necessitating high working-memory capacity to integrate scientific arguments, maintain academic tone, avoid grammatical anomalies, and ensure scholarly rigor. This subscale measures behaviors including:
- Automated drafting of manuscript and thesis paragraphs.
- Rephrasing, paraphrasing, and stylistic elevation of user-authored drafts into conventional scholarly English.
- Generating alternative rhetorical perspectives and constructing formal scientific counterarguments.
- Streamlining technical sentence structures and enhancing text cohesion.
2. Academic Task Support
The Academic Task Support dimension encompasses the deployment of generative AI across peripheral and structural phases of the academic workflow. Rather than directly generating publishable text, students utilize the conversational agent as a meta-cognitive collaborator, organizational mechanism, and exploratory thinking partner. This subscale evaluates behaviors such as:
- Alleviating “writer’s block” through open-ended conceptual brainstorming.
- Structuring and organizing complex thematic thoughts, lecture notes, or literature matrices into structured outlines.
- Conducting preliminary scoping of unfamiliar academic domains to locate relevant concepts.
- Synthesizing instructional materials, self-quizzing modules, and study summaries to optimize personal exam preparation and research workflow efficiency.
3. Reliance and Trust
The Reliance and Trust dimension taps into the psychological, affective, and epistemic relationship between the student and the conversational agent. This dimension evaluates the degree of subjective dependence, attribution of epistemic authority, and uncritical acceptance of AI-generated responses. Sub-facets captured within this domain include:
- Epistemic trust in the factual accuracy, analytical competence, and objectivity of ChatGPT’s outputs.
- Habitual dependence on the conversational agent as the default arbiter of conceptual verification, iterative feedback, and brainstorming.
- The perceived indispensability of AI tools for sustaining normative academic productivity, illustrating high cognitive vulnerability should access to the tool be revoked.
Theoretical Framework
The conceptual architecture of the ChatGPT Usage Scale is rooted in three complementary psychological and educational frameworks: the Technology Acceptance Model, Cognitive Load Theory, and Self-Determination Theory.
Technology Acceptance Model (TAM)
Originating from the work of Fred Davis (1989) and subsequently expanded by Venkatesh and Davis (2000), the Technology Acceptance Model posits that the actual utilization of an emerging information system is determined by a behavioral intention to use, which is jointly shaped by two primary cognitive appraisals: Perceived Usefulness (PU)—the degree to which an individual believes that using a specific system will enhance their job or academic performance—and Perceived Ease of Use (PEOU)—the degree to which an individual expects the target system to be free of cognitive or physical effort.
In the context of the ChatGPT Usage Scale, the TAM framework explains why postgraduate students readily integrate conversational LLMs into their scholarly routines. Unlike legacy academic database querying or traditional search engine optimization, conversational AI presents a frictionless natural-language interface (high PEOU) coupled with immediate, synthetically integrated text responses that yield immediate academic productivity gains (high PU). The subscales Academic Writing Aid and Academic Task Support represent direct manifestations of operationalized Perceived Usefulness, measuring the perceived and realized utility of AI in compressing the time required to complete scholarly milestones.
Cognitive Load Theory (CLT)
Developed by John Sweller (1988) and extended by Chen, Kalyuga, and Sweller (2015), Cognitive Load Theory operates on the universal architecture of human cognition: working memory is strictly finite in both capacity and duration, whereas long-term memory is functionally limitless. CLT differentiates among three varieties of cognitive load:
- Intrinsic load: The inherent difficulty associated with the instructional material or complex research problem itself.
- Extraneous load: Mental effort wasted due to the suboptimal presentation, organization, or mechanics of the task.
- Germane load: Mental processing dedicated to schema acquisition, deep structural comprehension, and conceptual mastery.
Within this theoretical lens, the ChatGPT Usage Scale measures mechanisms of cognitive offloading. By offloading extraneous syntactic demands, preliminary drafting mechanics, and data organization onto the AI model, students theoretically conserve working-memory bandwidth. This freed mental capacity can then be redirected toward high-level theoretical synthesis and creative analysis (germane load). However, CLT also provides a critical warning: if offloading is applied indiscriminately to core analytical tasks, it eliminates the necessary “desirable difficulties” required for deep conceptual learning, producing superficial knowledge acquisition and heavy reliance.
Self-Determination Theory (SDT)
Formulated by Edward Deci and Richard Ryan (1985), Self-Determination Theory articulates that human flourishing and intrinsic motivation are predicated on the satisfaction of three fundamental psychological needs: competence, autonomy, and relatedness. In academic contexts, students strive to experience efficacy in challenging tasks and to feel in control of their educational trajectories. The Reliance and Trust subscale measures how students resolve these motivational drives in an AI-mediated environment: while conversational agents can temporarily inflate a student’s subjective sense of competence through rapid generation of sophisticated prose, excessive reliance may subtly erode intrinsic academic autonomy, binding the student’s sense of competence to external machine outputs.
Validity
The psychometric validation of the ChatGPT Usage Scale was executed through a rigorous empirical design evaluating structural, convergent, and construct validity within a large postgraduate cohort.
Structural Validity
The internal structural validity of the scale was established through sequential factor analytic phases. Following initial exploratory factor extraction, confirmatory factor analysis (CFA) evaluated whether the observed empirical covariance matrix conformed to the hypothesized three-factor multidimensional model. The CFA model demonstrated statistically significant standardized factor loadings across all 15 retained indicators, with loadings ranging from a moderate 0.434 to a robust 0.728. Every standardized path coefficient reached statistical significance ($p < 0.001$), confirming that each individual indicator serves as a valid empirical manifestation of its respective latent factor. The structural validity of a second-order factor model was also substantiated, confirming that the three first-order factors load reliably onto a generalized higher-order latent construct representing global academic ChatGPT usage.
Convergent Validity
Convergent validity demonstrates that items hypothesized to measure a single theoretical construct share a high proportion of common variance. In psychometric structural equation modeling, this is empirically tested using the Fornell and Larcker (1981) criterion via the Average Variance Extracted (AVE) statistic:
$$\text{AVE} = \frac{\sum_{i=1}^{k} \lambda_i^2}{k}$$
Where $\lambda_i$ represents the standardized factor loading of item $i$, and $k$ represents the number of indicators. The ChatGPT Usage Scale achieved an overall AVE value of 0.664. Because this figure substantially exceeds the widely accepted psychometric benchmark of 0.50, it provides decisive empirical evidence that more than 66% of the variance captured by the indicators is accounted for by the underlying latent factors, rather than residual measurement error. Furthermore, the Composite Reliability (CR) coefficient reached 0.855, far surpassing the standard 0.70 threshold for acceptable construct reliability.
Discriminant and Nomological Validity
The distinctiveness of the three extracted sub-dimensions was confirmed as inter-factor correlations remained within moderate boundaries ($r < 0.85$), ensuring that multicollinearity did not obscure the empirical separation among Academic Writing Aid, Academic Task Support, and Reliance and Trust. The scale demonstrated nomological validity by exhibiting predictable, statistically significant correlations with external educational constructs, aligning with prior academic literature on generative AI adoption and academic task engagement.
Reliability
Reliability refers to the precision, internal stability, and consistency of an instrument across repeated measurements. The ChatGPT Usage Scale was subjected to multi-tiered reliability diagnostics that moved beyond basic classical test theory limitations.
Internal Consistency: Cronbach’s Alpha
Under Classical Test Theory (CTT), internal consistency is conventionally indexed via Cronbach’s alpha (α). For the omnibus 15-item ChatGPT Usage Scale, the calculated overall coefficient was:
$$\alpha = 0.848 \approx 0.85$$
In psychometrics (Nunnally & Bernstein, 1994), an alpha value spanning between 0.80 and 0.90 is considered optimal for psychological and educational measurement tools. It confirms that the items reflect a cohesive construct without introducing excessive redundancy or bloating the item pool with near-identical paraphrasing.
Tau-Equivalence and McDonald’s Omega
A recognized limitation of Cronbach’s alpha is its reliance on the assumption of essential tau-equivalence—the presupposition that all items measure the latent construct with equal precision and possess identical factor loadings. When tau-equivalence is violated (as is typical in multidimensional or unequal-loading behavioral scales), Cronbach’s alpha systematically underestimates true reliability. To provide a more robust estimation, the authors computed McDonald’s omega (ω):
$$\omega = 0.849 \approx 0.85$$
The alignment between McDonald’s omega ($\omega = 0.849$) and Cronbach’s alpha ($lpha = 0.848$), along with a Composite Reliability of 0.855, confirms that the scale possesses high internal stability and measurement precision across diverse respondent profiles.
Factor Analysis
The factorial validity and latent dimensionality of the ChatGPT Usage Scale were derived using a comprehensive two-stage factor analytic pipeline: an initial Exploratory Factor Analysis (EFA) followed by Confirmatory Factor Analysis (CFA).
Exploratory Factor Analysis (EFA)
The original conceptual pool assembled by the researchers consisted of 39 candidate items intended to capture every facet of AI interaction in higher education. The EFA was conducted using Principal Component Analysis (PCA) accompanied by orthogonal Varimax rotation to maximize the variance of the squared loadings across factors, thereby generating an interpretable simple structure. The retention of items was governed by strict psychometric criteria:
- Retention of factors displaying eigenvalues > 1.0 (Kaiser-Guttman rule) coupled with scree plot inspection.
- Suppression and elimination of items with primary factor loadings < 0.40.
- Elimination of complex cross-loading items (items exhibiting cross-loadings ≥ 0.35 on two or more factors).
Through this iterative refinement process, 24 suboptimal items were eliminated, resulting in a 15-item structure spanning three distinct factors that cumulatively accounted for approximately 49% of the total variance.
Confirmatory Factor Analysis (CFA)
To confirm the latent architecture derived from the EFA, a Confirmatory Factor Analysis was estimated using structural equation modeling software. The goodness-of-fit of the hypothesized measurement model was evaluated against conventional structural benchmarks established by Hu and Bentler (1999) and Hooper, Coughlan, and Mullen (2008):
| Fit Index | Observed Value | Conventional Threshold | Interpretation |
|---|---|---|---|
| CFI (Comparative Fit Index) | 0.917 | ≥ 0.90 (Acceptable) / ≥ 0.95 (Good) | Acceptable Model Fit |
| TLI (Tucker-Lewis Index) | 0.900 | ≥ 0.90 (Acceptable) / ≥ 0.95 (Good) | Acceptable Model Fit |
| RMSEA (Root Mean Square Error of Approx.) | 0.060 | ≤ 0.08 (Acceptable) / ≤ 0.05 (Close Fit) | Good/Acceptable Fit |
The confirmatory fit indices demonstrated that the three-factor model adequately captures the empirical data structure. Standardized factor loadings across all 15 indicators ranged from 0.434 to 0.728, providing structural evidence supporting the tripartite conceptualization of academic AI use.
Instrument / Measurement Tool
- Test Type: Psychometric self-report questionnaire / Behavioral rating inventory.
- Target Population: Postgraduate students (enrolled in Postgraduate Diploma, Master of Science/Arts, or Doctor of Philosophy/Education programs) and advanced higher education scholars.
- Standardization Sample: 443 postgraduate scholars (194 male, 249 female; mean age = 27.4 years, SD = 4.8) from Kafr el-Sheikh University and Al-Azhar University, Egypt.
- Total Item Count: 15 items (distilled from an initial 39-item experimental pool).
- Latent Factor Structure: 3-factor multidimensional model:
- Factor 1: Academic Writing Aid (6 items: original item indices 34, 38, 13, 5, 18, 23).
- Factor 2: Academic Task Support (5 items: original item indices 8, 30, 14, 3, 10).
- Factor 3: Reliance and Trust (4 items: original item indices 12, 2, 16, 22).
- Response Scale: 15 items, 5-point Likert scale (1 = strongly disagree to 5 = strongly agree):
- 1 = Strongly Disagree
- 2 = Disagree
- 3 = Neutral / Undecided
- 4 = Agree
- 5 = Strongly Agree
- Scoring Protocol: Subscale and global composite scores are generated by either calculating the direct sum or computing the arithmetic mean across indicators. Higher composite scores correspond to higher frequency, reliance, and functional integration of ChatGPT in academic activities.
- Administration Modality: Self-administered online survey or paper-pencil test; estimated completion time is approximately 5 to 7 minutes.
Permissions & Fee and Test Year
- Year of Initial Publication: 2024.
- Primary Publication: Published in BMC Psychology (BioMed Central / Springer Nature). DOI: 10.1186/s40359-024-01983-4.
- Permissions and Access: The academic article is distributed under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted non-commercial and commercial academic use, reproduction, and distribution in any medium, provided appropriate credit is given to the original authors.
- Accessing Item Text: Because individual item statements were not fully reproduced in the open-access text, researchers wishing to deploy the authorized, exact questionnaire items for research or institutional assessment should directly contact the corresponding author, Mohamed Nemt-Allah, via email ([email protected]).
- Licensing Fee: No commercial license fees or financial payments are required for non-profit academic research, university institutional evaluation, or thesis investigations.
References
Abdaljaleel, M., Barakat, M., Alsanafi, M., Salim, N., Abazid, H., Malaeb, D., & Sallam, M. (2023). Factors influencing attitudes of university students towards ChatGPT and its usage: A multi-national study validating the TAME-ChatGPT survey instrument. Preprints 2023, 2023090541. https://doi.org/10.20944/preprints202309.1541.v1
Aydin, Ö., & Karaarslan, E. (2023). Is ChatGPT leading generative AI? What is beyond expectations? Academic Platform Journal of Engineering and Smart Systems, 11(3), 118–134. https://doi.org/10.21541/apjess.1293702
Bin-Nashwan, S. A., Sadallah, M., & Bouteraa, M. (2023). Use of ChatGPT in academia: Academic integrity hangs in the balance. Technology in Society, 75, Article 102370. https://doi.org/10.1016/j.techsoc.2023.102370
Chen, O., Kalyuga, S., & Sweller, J. (2015). The worked example effect, the generation effect, and element interactivity. Journal of Educational Psychology, 107(3), 689–704. https://doi.org/10.1037/edu0000018
Davis, F. D. (1989). Perceived usefulness, perceived ease of use, and user acceptance of information technology. MIS Quarterly, 13(3), 319–340. https://doi.org/10.2307/249008
Deci, E. L., & Ryan, R. M. (1985). The general causality orientations scale: Self-determination in personality. Journal of Research in Personality, 19(2), 109–134. https://doi.org/10.1016/0092-6566(85)90023-6
Elbably, Y., & Nemt-Allah, M. (2024). Grand challenges for ChatGPT usage in education: Psychological theories, perspectives and opportunities. Psychological Research in Education and Social Sciences, 5(2), 31–36.
Floridi, L., & Chiriatti, M. (2020). GPT-3: Its nature, scope, limits, and consequences. Minds and Machines, 30(4), 681–694. https://doi.org/10.1007/s11023-020-09548-1
Fornell, C., & Larcker, D. F. (1981). Evaluating structural equation models with unobservable variables and measurement error. Journal of Marketing Research, 18(1), 39–50. https://doi.org/10.1177/002224378101800104
Hair, J. F., Black, W. C., Babin, B. J., & Anderson, R. E. (2014). Multivariate data analysis (7th ed.). Pearson Education Limited.
Hartley, K., Hayak, M., & Ko, U. (2024). Artificial intelligence supporting independent student learning: An evaluative case study of ChatGPT and learning to code. Education Sciences, 14(2), Article 120. https://doi.org/10.3390/educsci14020120
Henderson, M., Finger, G., & Selwyn, N. (2016). What’s used and what’s useful? Exploring digital technology use(s) among taught postgraduate students. Active Learning in Higher Education, 17(3), 235–247. https://doi.org/10.1177/1469787416654798
Hooper, D., Coughlan, J., & Mullen, M. (2008). Structural equation modelling: Guidelines for determining model fit. Electronic Journal of Business Research Methods, 6(1), 53–60.
Hu, L., & Bentler, P. M. (1999). Cutoff criteria for fit indexes in covariance structure analysis: Conventional criteria versus new alternatives. Structural Equation Modeling: A Multidisciplinary Journal, 6(1), 1–55. https://doi.org/10.1080/10705519909540118
Huallpa, J. (2023). Exploring the ethical considerations of using ChatGPT in university education. Periodicals of Engineering and Natural Sciences, 11(4), 105–115.
İpek, Z., Gözüm, A., Papadakis, S., & Kallogiannakis, M. (2023). Educational applications of the ChatGPT AI system: A systematic review research. Educational Process: International Journal, 12(3), 26–55. https://doi.org/10.22521/edupij.2023.123.2
Nemt-Allah, M., Khalifa, W., Badawy, M., Elbably, Y., & Ibrahim, A. (2024). ChatGPT Usage Scale. BMC Psychology. https://doi.org/10.1186/s40359-024-01983-4
Ng, J. Y., Ntoumanis, N., Thøgersen-Ntoumani, C., Deci, E. L., Ryan, R. M., Duda, J. L., & Williams, G. C. (2012). Self-determination theory applied to health contexts: A meta-analysis. Perspectives on Psychological Science, 7(4), 325–340. https://doi.org/10.1177/1745691612447309
Nunnally, J. C., & Bernstein, I. H. (1994). Psychometric theory (3rd ed.). McGraw-Hill.
Sain, Z. H., & Hebebci, M. T. (2023). ChatGPT and beyond: The rise of AI assistants and chatbots in higher education. In S. M. Curle & M. T. Hebebci (Eds.), Proceedings of International Conference on Academic Studies in Technology and Education 2023 (pp. 1–12). ARSTE Organization.
Sallam, M., Salim, N., Barakat, M., Al-Mahzoum, K., Ala’a, B., Malaeb, D., & Hallit, S. (2023). Assessing health students’ attitudes and usage of ChatGPT in Jordan: Validation study. JMIR Medical Education, 9(1), Article e48254. https://doi.org/10.2196/48254
Schön, E. M., Neumann, M., Hofmann-Stölting, C., Baeza-Yates, R., & Rauschenberger, M. (2023). How are AI assistants changing higher education? Frontiers in Computer Science, 5, Article 1208550. https://doi.org/10.3389/fcomp.2023.1208550
Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257–285. https://doi.org/10.1207/s15516709cog1202_4
Venkatesh, V., & Davis, F. D. (2000). A theoretical extension of the Technology Acceptance Model: Four longitudinal field studies. Management Science, 46(2), 186–204. https://doi.org/10.1287/mnsc.46.2.186.11926
Wang, T., Díaz, D. V., Brown, C., & Chen, Y. (2023). Exploring the role of AI assistants in computer science education: Methods, implications, and instructor perspectives. In 2023 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC) (pp. 92–102). IEEE. https://doi.org/10.1109/VL-HCC57772.2023.00018
Worthington, R. L., & Whittaker, T. A. (2006). Scale development research: A content analysis and recommendations for best practices. The Counseling Psychologist, 34(6), 806–838. https://doi.org/10.1177/0011000006288127
Zawacki-Richter, O., Marín, V. I., Bond, M., & Gouverneur, F. (2019). Systematic review of research on artificial intelligence applications in higher education—where are the educators? International Journal of Educational Technology in Higher Education, 16(1), Article 27. https://doi.org/10.1186/s41239-019-0171-0
Items of the Scale
The individual item statements of the ChatGPT Usage Scale are not published in full in the open public domain. To maintain measurement fidelity and support ongoing psychometric standardization, the original developers (Mohamed Nemt-Allah et al., 2024) retain proprietary oversight of the questionnaire inventory. Researchers, academic departments, and clinicians requiring access to the full, authorized 15-item inventory must contact the corresponding author directly.
Instrument Architecture & Subscale Mapping
The scale consists of 15 validated items structured into three distinct sub-domains, identified from the original 39-item development pool:
- Subscale 1: Academic Writing Aid (6 Items)
- Assigned Item Numbers: 34, 38, 13, 5, 18, and 23.
- Functional Scope: Evaluates the extent to which scholars deploy ChatGPT for direct textual composition, generating draft prose, rephrasing technical formulations, stylistic refining, and synthesizing formal scholarly counterarguments.
- Subscale 2: Academic Task Support (5 Items)
- Assigned Item Numbers: 8, 30, 14, 3, and 10.
- Functional Scope: Assesses the utilization of ChatGPT for non-compositional research and study workflows, such as overcoming creative inertia (“writer’s block”), conceptual planning, organizing thematic notes, preliminary domain scoping, and structuring revision materials.
- Subscale 3: Reliance and Trust (4 Items)
- Assigned Item Numbers: 12, 2, 16, and 22.
- Functional Scope: Measures the student’s internal cognitive and affective dependence on conversational AI, epistemic trust in the veracity and objectivity of generated outputs, and perceived inability to sustain research productivity without continuous AI feedback.
Mandatory Response Format
All items on the instrument are rated using a uniform 15 items, 5-point Likert scale (1 = strongly disagree to 5 = strongly agree):
Scoring and Computation Formula
In accordance with the scoring instructions specified by Nemt-Allah et al. (2024), scores can be calculated either by summing or averaging item responses:
- Subscale Scores: Computed by calculating the mean or sum of the items corresponding to that specific subscale:
- Academic Writing Aid: Sum or mean of items 34, 38, 13, 5, 18, 23 (Score range: 6 to 30 for summation; 1.0 to 5.0 for mean).
- Academic Task Support: Sum or mean of items 8, 30, 14, 3, 10 (Score range: 5 to 25 for summation; 1.0 to 5.0 for mean).
- Reliance and Trust: Sum or mean of items 12, 2, 16, 22 (Score range: 4 to 20 for summation; 1.0 to 5.0 for mean).
- Total Composite Score: Calculated by summing all 15 items (score range: 15 to 75) or averaging all 15 items (score range: 1.0 to 5.0). Higher values indicate greater academic integration, functional reliance, and cognitive trust in ChatGPT.