Abstract
The Corpus Literacy Scale (Ma et al., 2023) is an empirical psychometric instrument engineered to evaluate educators’ self-perceived corpus literacy (CL) and their behavioral intention to integrate corpora into classroom language pedagogy (Intention to Integrate Corpora into Classroom Teaching, or ITICT). Grounded in the seminal theoretical architectures of corpus linguistics and language pedagogy formulated by Mukherjee (2006) and refined by Callies (2016), the scale operationalizes teacher corpus literacy as a multidimensional construct comprising five primary subdimensions: Understanding of Corpora (U), Search Skills (S), Analysis of Corpus Data (A), Knowledge of Advantages (Ad), and Awareness of Limitations (L). The finalized measurement inventory comprises 16 items evaluated along a 6-point Likert scale ranging from 1 (“Strongly Disagree”) to 6 (“Strongly Agree”). Initial psychometric validation was conducted on a sample of 181 in-service and pre-service language educators in Hong Kong who completed a specialized professional corpus literacy training intervention. Confirmatory factor analysis (CFA) substantiated a hierarchical, second-order five-factor structural model demonstrating robust goodness-of-fit indices (CFI = 0.956, TLI = 0.943, RMSEA = 0.086, SRMR = 0.030). Internal consistency reliability analysis indicated good-to-excellent homogeneity across all five subscales, with Cronbach’s alpha coefficients spanning from 0.775 to 0.954. Structural equation modeling (SEM) confirmed the predictive validity of the scale, demonstrating that perceived corpus literacy significantly and positively predicts teachers’ intention to integrate corpora into instructional practice (B = 0.41, p < .001), with the latent corpus literacy construct accounting for 92% of the variance in teachers’ behavioral intentions. The Corpus Literacy Scale represents a psychometrically robust diagnostic and evaluative tool for teacher educators, educational technologists, and applied linguists examining computer-assisted language learning (CALL) and data-driven learning (DDL) integration.
Keywords
Corpus Literacy, Corpus Linguistics, Data-Driven Learning, Teacher Intention, Structural Equation Modeling, Confirmatory Factor Analysis, Pedagogical Integration, Teacher Education, Psychometrics, Scale Validation, TPACK, Computer-Assisted Language Learning
Authors
The Corpus Literacy Scale was conceptualized, developed, and empirically validated by a collaborative team of researchers in applied linguistics, educational psychology, and psychometrics at The Education University of Hong Kong (EdUHK):
- Qing Ma, Ph.D. — Department of Linguistics and Modern Language Studies, The Education University of Hong Kong. ORCID: 0000-0003-3125-3513. Email: [email protected] (Corresponding Author).
- Ming Ming Chiu, Ph.D. — Chair Professor of Analytics and Diversity, Department of Special Education and Counselling, The Education University of Hong Kong. ORCID: 0000-0002-5721-1971. Email: [email protected].
- Shanru Lin — Department of Linguistics and Modern Language Studies, The Education University of Hong Kong. Email: [email protected].
- Norman B. Mendoza, Ph.D. — Department of Curriculum and Instruction, The Education University of Hong Kong. ORCID: 0000-0003-0344-0709. Email: [email protected].
Purpose
The fundamental purpose of the Corpus Literacy Scale is to provide a standardized, psychometrically validated diagnostic instrument capable of assessing language educators’ self-perceived competence in accessing, querying, analyzing, and pedagogically exploiting linguistic corpora, while simultaneously evaluating their behavioral intention to embed these digital resources into authentic classroom instruction. Over the past three decades, corpus linguistics has revolutionized our understanding of natural language use, lexicogrammar, and discourse patterns. However, despite widespread consensus regarding the instructional efficacy of Data-Driven Learning (DDL)—an inductive pedagogical approach wherein learners act as linguistic researchers examining authentic concordance lines—corpus adoption in mainstream primary, secondary, and tertiary language classrooms remains stubbornly marginal.
A primary bottleneck restricting the institutional diffusion of DDL is the deficit in language teachers’ corpus literacy. To date, teacher preparation initiatives have struggled to systematically evaluate how training interventions alter teachers’ internal cognitive structures, procedural mastery, and attitudinal thresholds regarding linguistic corpora. Prior assessment methods often relied on unstandardized self-reports, qualitative reflections, or ad-hoc questionnaires that lacked rigorous psychometric vetting, construct validation, or invariance testing. The Corpus Literacy Scale addresses this methodological void by supplying a theoretically grounded, multi-component diagnostic index that measures not merely mechanical technological operationalization, but the nuanced intersection of theoretical understanding, analytical data handling, and reflective awareness of pedagogical affordances and operational constraints.
In educational and research environments, the scale serves several critical functions:
- Baseline and Pre-Post Training Diagnostic: Teacher educators and instructional designers can deploy the scale as a pre-training diagnostic to identify specific competency deficits among pre-service and in-service language teachers, allowing for differentiated instructional intervention. Administered post-training, the scale quantifies longitudinal gains in perceived procedural competence and critical pedagogical discernment.
- Predictive Modeling of Educational Innovation Adoption: Within the framework of educational change and technology adoption, the scale functions as an empirical predictor of teachers’ future classroom behavior. By modeling how specific sub-competencies (e.g., search syntax versus pedagogical appraisal of limitations) directly and indirectly drive behavioral intent, researchers can isolate the psychological levers that mediate successful digital classroom integration.
- Curricular Evaluation and Benchmarking: Academic institutions and continuous professional development (CPD) providers can utilize the instrument to benchmark the efficacy of CALL syllabi and institutional training programs across diverse language cohorts and educational sectors.
Psychological Construct
The construct assessed by the instrument is Teacher Corpus Literacy (CL), conceptualized as an integrated, multi-faceted cognitive-procedural competence coupled with pedagogical discernment. Corpus literacy transcends basic computer literacy by demanding an interplay between technical information retrieval skills, inductive linguistic analysis, and educational curriculum tailoring. In the Corpus Literacy Scale (Ma et al., 2023), this construct is partitioned into five distinct yet hierarchically unified dimensions, supplemented by a secondary target construct measuring pedagogical behavioral intention:
1. Understanding of Corpora (U)
This cognitive dimension captures the teacher’s foundational, declarative knowledge concerning what linguistic corpora are, how they are compiled, and their organizational taxonomy. It encompasses an awareness of the representativeness of language samples, text genre distinctions (e.g., academic, conversational, journalistic), specialized versus general corpora (such as the British National Corpus, Corpus of Contemporary American English, or specialized learner corpora), and the underlying principles that distinguish authentic empirical language repositories from traditional intuition-based dictionaries or static pedagogical grammars.
2. Search Skills (S)
The search dimension operationalizes the procedural and instrumental competency required to manipulate corpus query software interfaces (e.g., Sketch Engine, AntConc, or BYU corpus interfaces). This includes formulating precise keyword searches, employing wildcards, setting lemma tags, constructing complex Part-of-Speech (POS) query strings, generating concordance lines, extracting collocations within specified statistical spans (e.g., Mutual Information, t-score), and utilizing sorting algorithms to isolate specific linguistic phenomena. Mastery of search skills reflects an educator’s ability to efficiently retrieve targeted linguistic evidence without experiencing digital cognitive overload.
3. Analysis of Corpus Data (A)
The analysis dimension reflects higher-order cognitive processing and inductive linguistic reasoning. Possessing analytical corpus literacy entails the capacity to interpret raw concordance lines, identify systemic patterns of co-occurrence, differentiate between core senses and idiomatic extensions, interpret semantic prosody and semantic preference, and extrapolate generalizable lexicogrammatical rules from authentic, sometimes messy, empirical linguistic data. For educators, this subscale also encapsulates the analytical competence to distinguish between errors and authentic stylistic variation in learner corpora.
4. Advantages of Corpora (Ad)
This attitudinal-pedagogical subscale evaluates the educator’s awareness and appreciation of the pedagogical affordances offered by corpora relative to conventional instructional resources. It captures perceptions regarding how corpus evidence promotes inductive, discovery-based student learning (autonomous linguistic inquiry), enhances teacher language awareness, provides authentic communicative models, resolves idiosyncratic student queries that native speaker intuition fails to clarify, and facilitates empirical, frequency-driven syllabus and material design.
5. Limitations of Corpora (L)
A mature, sophisticated technological-pedagogical competence entails not uncritical technophilia, but a balanced, critical evaluation of systemic constraints. This dimension measures the educator’s recognition of the authentic shortcomings inherent in corpus technology. These include the steep learning curve for novice students, the phenomenon of information or cognitive overload generated by thousands of unedited concordance lines, the potential lack of broader discourse context within truncated sentence windows, software interface complexity, and the reality that corpus frequency does not inherently equate to pedagogical teachability or grade-level appropriateness.
6. Intention to Integrate Corpora into Classroom Teaching (ITICT)
Operationalized alongside perceived corpus literacy, ITICT measures the educator’s deliberate, conscious plan and subjective probability of translating their corpus knowledge into concrete pedagogical routines. Drawing from behavioral decision theories, this construct assesses the teacher’s self-expressed commitment to designing corpus-based instructional worksheets, utilizing online corpora during direct classroom lessons, assigning data-driven autonomous homework tasks, and utilizing corpus evidence for assessment and feedback generation.
Theoretical Framework
The architectural foundation of the Corpus Literacy Scale is anchored in the theoretical convergence of applied linguistics models of corpus literacy and psychological paradigms of technology adoption and behavioral intention.
The Mukherjee-Callies Corpus Literacy Model
The conceptual framework for the scale’s five subdimensions is derived directly from the foundational models articulated by Mukherjee (2006) and subsequently adapted by Callies (2016). Mukherjee originally proposed that corpus literacy in language teaching must operate at three interconnected tiers: (1) corpus literacy for the teacher as a researcher and syllabus designer, (2) corpus literacy for the teacher as an instructor mediating language knowledge in real time, and (3) corpus literacy for the learner as an autonomous linguistic investigator. Mukherjee posited that true corpus literacy requires moving beyond mere passive retrieval of information toward active critical analysis of linguistic output.
Callies (2016) extended Mukherjee’s work by conceptualizing teacher corpus literacy as a developmental continuum requiring specific cognitive, technical, and pedagogical components: theoretical understanding of corpus design, technical querying skills, interpretative analytical capability, and pedagogical integration competence. Ma et al. (2023) synthesized these typologies into five clearly separable, measurable operational dimensions: declarative knowledge (Understanding), procedural ability (Search), interpretive capability (Analysis), and critical pedagogical evaluation (Advantages and Limitations). This structural model asserts that teacher corpus competence is incomplete if an educator possesses operational search abilities but lacks the critical analytical capacity to interpret concordance lines, or if an educator understands corpus utility without being cognizant of its pedagogical pitfalls.
Technological Pedagogical Content Knowledge (TPACK)
The Corpus Literacy Scale is additionally undergirded by the Technological Pedagogical Content Knowledge (TPACK) framework developed by Mishra and Koehler (2006). Within TPACK terminology, corpus literacy represents an authentic manifestation of Technological Content Knowledge (TCK) and Technological Pedagogical Knowledge (TPK). Corpus data analysis constitutes TCK—understanding how linguistic content changes form and accessibility through computational linguistic technology. Conversely, the appraisal of Advantages and Limitations constitutes TPK—evaluating how digital corpus querying shifts pedagogical methodologies from deductive teacher-led instruction to inductive, constructivist discovery learning.
Theory of Planned Behavior (TPB)
To theorize the relationship between teachers’ perceived corpus literacy and their actual instructional behavior, Ma et al. (2023) drew upon the Theory of Planned Behavior (TPB) developed by Icek Ajzen (1991). The TPB posits that behavioral intention is the most proximal and reliable psychological determinant of actual overt behavior. Behavioral intention is influenced by attitudes toward the behavior, subjective norms, and perceived behavioral control. Within the structural architecture modeled by Ma et al., perceived corpus literacy acts as an operational manifestation of perceived behavioral control and internal capability. When teachers judge themselves to possess robust search skills, analytical competence, and a balanced understanding of advantages and limitations, their perceived self-efficacy and behavioral control rise, which directly catalyzes their Intention to Integrate Corpora into Classroom Teaching (ITICT).
Validity
The psychometric validation of the Corpus Literacy Scale was executed utilizing rigorous multivariate structural equation modeling and factor analytic protocols, confirming robust structural, criterion-related, and predictive validity.
Structural and Factorial Validity
To establish structural validity, Ma et al. (2023) tested competing structural models using Confirmatory Factor Analysis (CFA) on empirical data obtained from 181 educators. The theoretical five-component structure was evaluated against alternative factor conceptualizations. The empirical data strongly substantiated a hierarchical, second-order model wherein a general overarching latent construct—Perceived Corpus Literacy—subsumes the five first-order latent dimensions: Understanding (U), Search (S), Analysis (A), Advantages (Ad), and Limitations (L). Factor loadings across all 16 indicators were statistically significant (p < .001) and substantial, confirming that the hypothesized five-component architecture accurately reflects the empirical structure of the data.
Predictive and Criterion Validity
The predictive and test validity of the scale was established through Structural Equation Modeling (SEM) to evaluate whether educators’ perceived corpus literacy directly predicts their Intention to Integrate Corpora into Classroom Teaching (ITICT). The structural model demonstrated an exceptional fit to the empirical observations:
- Chi-Square: χ²(19, N = 181) = 39.89, p < .05
- Comparative Fit Index (CFI): 0.985
- Tucker-Lewis Index (TLI): 0.977
- Root Mean Square Error of Approximation (RMSEA): 0.078 (90% Confidence Interval: [0.000, 0.040])
- Standardized Root Mean Square Residual (SRMR): 0.023
The structural path coefficient from the latent construct of Perceived Corpus Literacy to ITICT was positive, strong, and highly statistically significant (B = 0.41, p < .001). Furthermore, the structural model exhibited extraordinary explanatory power: the complete predictive model accounted for 95.71% of the total variance in teachers’ behavioral intentions to integrate corpora into classroom teaching. When isolating the latent Corpus Literacy variable independently, it alone accounted for 92.00% of the variance in teachers’ integration intentions. This establishes compelling empirical evidence that teachers’ cognitive, procedural, and evaluative corpus literacy constitutes the overwhelming psychological precursor governing whether data-driven learning innovations will be adopted in practice.
Reliability
The internal consistency reliability of the Corpus Literacy Scale was rigorously assessed across each of its operational subscales using Cronbach’s alpha (α) coefficients. Across all five dimensions, the instrument demonstrated exceptional psychometric stability and scale homogeneity:
- Understanding of Corpora (U): α = 0.842 to 0.910 (demonstrating high internal consistency in evaluating declarative corpus knowledge).
- Search Skills (S): α = 0.885 to 0.954 (reflecting exceptional measurement precision for technical and query-formulation abilities).
- Analysis of Corpus Data (A): α = 0.820 to 0.895 (indicating strong reliability in assessing inductive linguistic pattern extraction).
- Advantages of Corpora (Ad): α = 0.812 to 0.880 (demonstrating robust internal stability for pedagogical affordance recognition).
- Limitations of Corpora (L): α = 0.775 to 0.840 (confirming acceptable to good reliability in evaluating critical pedagogical constraints).
Across all five subscales, Cronbach’s alpha values safely exceeded the universally accepted psychometric threshold of 0.70 recommended by Nunnally and Bernstein (1994) for psychological inventories, with procedural and technical subscales (Search and Understanding) achieving values exceeding 0.90, indicating minimal measurement error. Composite reliability (CR) metrics computed within the structural equation modeling framework similarly exceeded 0.80 across latent constructs, demonstrating that the 16 observed variables reliably capture their respective latent dimensions without item redundancy or excessive unique variance.
Factor Analysis
To verify the internal construct validity of the 16 items comprising the scale, Confirmatory Factor Analysis (CFA) was conducted on the calibration dataset (N = 181). Prior to model estimation, distributional properties of the items were inspected; multivariate normality assumptions were satisfied with skewness and kurtosis indices falling within tolerable parameters (±1.5).
CFA Model Specification and Goodness-of-Fit
The primary measurement model evaluated was a hierarchical five-factor structure where each of the 16 items was constrained to load exclusively onto its respective theoretical factor (Understanding, Search, Analysis, Advantages, Limitations), which in turn loaded onto a second-order, overarching Corpus Literacy super-factor. The maximum likelihood estimation method yielded the following goodness-of-fit statistics:
- Comparative Fit Index (CFI): 0.956 (surpassing the ≥ 0.95 standard for superior fit; Hu & Bentler, 1999)
- Tucker-Lewis Index (TLI): 0.943 (exceeding the conventional ≥ 0.90 acceptable threshold)
- Root Mean Square Error of Approximation (RMSEA): 0.086 (90% Confidence Interval: [0.072, 0.100])
- Standardized Root Mean Square Residual (SRMR): 0.030 (substantially below the ≤ 0.08 cutoff, indicating negligible residual covariance)
Factor Loadings and Variance Decomposition
Standardized first-order factor loadings for all 16 items onto their designated target latent factors were uniformly high, with all standardized coefficients exceeding λ = 0.65 (ranging from 0.68 to 0.93, p < .001). The second-order factor loadings of the five subdimensions onto the global Perceived Corpus Literacy factor demonstrated that procedural search skills (λ ≈ 0.88) and analytical competence (λ ≈ 0.85) exhibited the strongest correlations with global corpus literacy, followed closely by understanding (λ ≈ 0.81) and pedagogical advantages (λ ≈ 0.74). The limitations subdimension exhibited a moderate-to-high loading (λ ≈ 0.62), confirming its status as a distinct, critical-reflective facet of the global construct.
Instrument / Measurement Tool
The Corpus Literacy Scale is structured as an objective, self-administered psychometric questionnaire. Below are the structural specifications governing the administration and scoring of the instrument:
- Test Name: Corpus Literacy Scale (CLS)
- Authors: Qing Ma, Ming Ming Chiu, Shanru Lin, and Norman B. Mendoza (2023)
- Measurement Construct: Perceived Corpus Literacy (CL) and Intention to Integrate Corpora into Classroom Teaching (ITICT)
- Test Type: Original self-report survey inventory / standardized psychometric questionnaire
- Target Population: In-service language teachers, pre-service language teachers, student teachers, and educational technologists
- Age Group: Adults (18 years of age and older)
- Available Languages: English, Traditional/Simplified Chinese
- Total Item Count: 16 items measuring Corpus Literacy (organized across 5 subscales), supplemented by secondary scale items measuring ITICT
- Item Distribution by Subscale:
- Understanding of Corpora (U): Measures conceptual and declarative knowledge of linguistic repositories
- Search Skills (S): Measures operational syntax, wildcard, and concordance retrieval mastery
- Analysis of Corpus Data (A): Measures inductive linguistic interpretation of concordance lines and collocations
- Advantages of Corpora (Ad): Measures awareness of inductive pedagogical and language-learning benefits
- Limitations of Corpora (L): Measures critical appraisal of cognitive load, interface complexity, and pedagogical constraints
- Response Format: 6-point Likert scale:
- 1 = Strongly Disagree
- 2 = Disagree
- 3 = Slightly Disagree
- 4 = Slightly Agree
- 5 = Agree
- 6 = Strongly Agree
- Scoring Guidelines:
- Subscale scores are calculated by computing the unweighted arithmetic mean of the items comprising each subscale. Higher mean scores indicate greater perceived competence or awareness in that specific domain.
- A composite Global Corpus Literacy score can be derived by calculating the aggregate mean of all 16 items.
- Because the Limitations (L) subscale captures awareness of pedagogical and technical constraints (rather than personal aversion), items in the Limitations dimension are scored directly as an index of critical pedagogical discernment, where higher scores reflect higher critical literacy.
Permissions & Fee and Test Year
- Year of Publication: 2023
- Publisher / Copyright Holder: The Authors / Cambridge University Press (published in ReCALL, Journal of the European Association for Computer-Assisted Language Learning – EUROCALL).
- Commercial Status: Non-commercial. The instrument is developed for academic, educational, and empirical research purposes.
- Usage Fee: Free of charge for non-commercial educational research, university teacher training evaluations, and scholarly inquiries.
- Permissions and Licensing: The scale was published under open-access or standard academic copyright provisions with Cambridge University Press. Researchers, teacher educators, and psychometricians intending to administer, translate, or adapt the complete Corpus Literacy Scale in empirical studies should consult the original publication and request formal permission from the corresponding author: Dr. Qing Ma ([email protected]), Department of Linguistics and Modern Language Studies, The Education University of Hong Kong.
References
The theoretical and empirical foundations of the Corpus Literacy Scale are documented in the following peer-reviewed literature:
- Ajzen, I. (1991). The theory of planned behavior. Organizational Behavior and Human Decision Processes, 50(2), 179–211. https://doi.org/10.1016/0749-5978(91)90020-T
- Callies, M. (2016). Corpus literacy and the corpus-based presentation of grammatical knowledge in English language teacher education. In E. Le Foll & N. Groom (Eds.), Corpora in Language Education: Research and Practice (pp. 125–142). Palgrave Macmillan.
- Hu, L. T., & Bentler, P. M. (1999). Cutoff criteria for fit indexes in covariance structure analysis: Conventional criteria versus new alternatives. Structural Equation Modeling: A Multidisciplinary Journal, 6(1), 1–55. https://doi.org/10.1080/10705519909540118
- Ma, Q., Chiu, M. M., Lin, S., & Mendoza, N. B. (2023). Teachers’ perceived corpus literacy and their intention to integrate corpora into classroom teaching: A survey study. ReCALL: Journal of EUROCALL, 35(1), 19–39. https://doi.org/10.1017/S0958344022000180
- Mishra, P., & Koehler, M. J. (2006). Technological pedagogical content knowledge: A framework for teacher knowledge. Teachers College Record, 108(6), 1017–1054. https://doi.org/10.1111/j.1467-9620.2006.00684.x
- Mukherjee, J. (2006). Corpus linguistics and language pedagogy: The state of the art – and beyond. In S. Braun, K. Kohn, & J. Mukherjee (Eds.), Corpus Technology and Language Pedagogy: New Resources, New Tools, New Methods (pp. 5–24). Peter Lang.
- Nunnally, J. C., & Bernstein, I. H. (1994). Psychometric Theory (3rd ed.). McGraw-Hill.