Educational PsychologyPsychometrics & DiagnosticsReading & Literacy Assessments

Analysis of Individualization Forms Test Battery

The Analysis of Individualization Forms Test Battery (AVI-toetspakket) is a standardized psychometric instrument designed to assess oral reading fluency, technical decoding proficiency, and reading developmental progression in elementary school children.

memjavad
PUBLISHED
Scientifically Reviewed · Dr. Marwa Abd-Alazim · September 12, 2026
Medically & Scientifically Reviewed Verified: September 12, 2026
Dr. Marwa Abd-Alazim Ph.D.
Professor of Psychology University of Kerbala
Review Criteria & Clinical Standards

This content undergoes rigorous scientific peer-review and medical editorial standards at Arab Psychology Network to ensure clinical accuracy, validity, and compliance with evidence-based guidelines from leading psychological and healthcare authorities (APA / WHO).

1. Abstract

The Analysis of Individualization Forms Test Battery (commonly known in Dutch educational psychometrics as the Analyse van Individualiseringsvormen toetspakket or AVI-toetspakket) is a standardized, criterion-referenced, and norm-referenced diagnostic instrument designed to assess oral reading fluency, technical decoding proficiency, and the developmental progression of text-level reading skills in primary school children. Developed principally by J. Visser, A. van Laarhoven, and A. ter Beek in conjunction with the National Institute for Educational Measurement (Cito) in 1994 and subsequently refined across multiple iterations, the instrument forms an integral component of the Dutch Student and Education Monitoring System (Leerling- en onderwijsvolgsysteem, or LOVS). The assessment battery evaluates the speed, accuracy, and automaticity with which an individual child decodes orthographically controlled connected prose of systematically ascending linguistic complexity.

The battery comprises 11 standardized graded reading passage cards, structured into parallel forms (Version A and Version B) to facilitate longitudinal monitoring and prevent practice or retest contamination. The instrument maps directly onto developmental educational stages ranging from mid-grade 3 (M3, corresponding to the Dutch educational designation equivalent to international Grade 1 / primary year 3) through end-of-grade 7 (E7, corresponding to late primary education) and an advanced ‘Plus’ level, with the baseline stage designated as AVI-Start. Test administration involves individualized oral reading observation wherein examiners record total reading latency alongside quantitative and qualitative decoding errors, including word substitutions, omissions, additions, and self-corrections. Psychometric evaluations demonstrate high internal consistency (parallel-form reliability coefficients ranging between r = .88 and .94) and robust criterion-related validity when juxtaposed with discrete word identification measures, such as the Three-Minute Test (Drie Minuten Test; DMT), standardized reading comprehension metrics, and teacher clinical judgments. Item Response Theory (IRT) modeling, specifically the Rasch calibration framework, underpins the vertical equating and scalar invariance across progressive text difficulty levels.

2. Keywords

Analysis of Individualization Forms Test Battery, AVI-toetspakket, oral reading fluency, technical reading, decoding accuracy, reading speed, psychometrics, Cito LOVS, Rasch modeling, Dutch orthography, dyslexia diagnosis, primary education assessment, criterion-referenced testing

3. Authors

The original standardized version of the Analysis of Individualization Forms Test Battery was developed by:

  • J. Visser — Psychometrician and Educational Researcher, Centraal Instituut voor Toetsontwikkeling (Cito), Arnhem, The Netherlands.
  • A. van Laarhoven — Specialist in Reading Diagnostics and Special Educational Needs, Centraal Instituut voor Toetsontwikkeling (Cito), Arnhem, The Netherlands.
  • A. ter Beek — Test Development Specialist, Primary Education Division, Centraal Instituut voor Toetsontwikkeling (Cito), Arnhem, The Netherlands.

Subsequent psychometric recalibrations and modernizations (notably the 2008 and 2018 revisions) were conducted by research teams at Cito, including F. Moelands, R. Kamphuis, and colleagues, in collaboration with the Expertisecentrum Nederlands (Radboud University Nijmegen).

4. Purpose

The primary purpose of the Analysis of Individualization Forms Test Battery is to deliver a standardized, ecologically valid, and diagnostically sensitive evaluation of technical reading competence in elementary school-aged children. Technical reading (technisch lezen) represents the cognitive-linguistic capacity to accurately, swiftly, and effortlessly translate printed orthographic symbols (graphemes) into their corresponding auditory-phonological representations (phonemes) and semantic constructs. Unlike isolated word-list decoding assessments, the battery captures the complex interplay between decoding mechanics, syntactic parsing, and contextual processing within connected discourse.

From a clinical and educational diagnostics standpoint, the instrument fulfills three cardinal objectives:

  1. Screening and Longitudinal Monitoring: Embedded within the national LOVS tracking architecture, the instrument is systematically administered at least twice annually (mid-academic year and end-of-academic year) across standard primary school cohorts. This systematic tracking enables educators and school psychologists to map a student’s developmental trajectory against established developmental benchmarks, facilitating early identification of reading deceleration, developmental delays, or anomalous trajectories indicative of specific learning disorders such as developmental dyslexia.
  2. Instructional Matching and Individualized Pedagogy: In accordance with its conceptual nomenclature—the analysis of forms of educational individualization—the battery explicitly guides differentiated classroom instruction. By determining a child’s instructional reading level (the level at which instructional scaffolding maximizes learning gains without inducing cognitive overload or learned helplessness) and independent reading level (the level suitable for autonomous recreational reading), educators can assign targeted reading materials that match the learner’s precise zone of proximal development.
  3. Clinical Evaluation and Multi-Tiered Support (RTI): In neurodevelopmental and child neuropsychological clinics, the battery acts as a gold-standard diagnostic metric within a Response to Intervention (RTI) framework. Persistent deficits across AVI cards—characterized by high error frequency, prolonged reading latency, or failed progression across successive normative evaluation windows—provide empirical data required for formal multidisciplinary dyslexia assessments, academic accommodations, and specialized clinical orthodidactic interventions.

5. Psychological Construct

The Analysis of Individualization Forms Test Battery operationalizes the multifaceted construct of oral reading fluency (ORF) and technical reading ability. Rather than treating reading as a unitary, unanalyzable process, the instrument’s psychometric architecture evaluates performance across two interrelated neurocognitive dimensions: decoding accuracy and decoding automaticity (speed).

Decoding Accuracy

Decoding accuracy reflects the fidelity of grapheme-to-phoneme conversion and orthographic word recognition. In the scoring rubric of the battery, accuracy is measured by the total frequency and nature of deviations from the target text. These deviations encompass:

  • Substitutions: Replacing a target word with another real word or pseudoword (e.g., reading “tree” as “three”, or substituting morphologically related forms like “runs” for “running”), indicating lexical competition or deficient orthographic precision.
  • Omissions: Skipping phonemes, syllables, or entire lexical tokens, reflecting visual-attentional slips, saccadic control anomalies, or rapid processing failures.
  • Additions: Inserting unprinted morphemes or words into the stream of speech, often reflecting unconstrained contextual guessing driven by poor bottom-up decoding.
  • Hesitations and Self-Corrections: Pauses exceeding standard phonetic transition thresholds or post-error corrections, which signify active metacognitive monitoring albeit hindered by sub-optimal lexical retrieval speeds.

Decoding Automaticity and Speed

Decoding automaticity is defined as the rapid, effortless translation of printed text into spoken output without the conscious allocation of central executive cognitive capacity. The battery measures this dimension via total reading duration (measured precisely in seconds) required to articulate each standardized passage. Automaticity is grounded in the psychological premise that working memory capacity is strictly bounded. When lower-level phonological decoding processes become automated, cognitive resources are liberated for higher-level discourse synthesis, syntactic chunking, and text comprehension.

Construct Interaction and Threshold Mastery

The diagnostic power of the battery lies in the convergence of accuracy and speed into discrete performance bands:

  • Beheersing (Mastery): The student reads the passage within the designated speed ceiling while committing fewer than the allowable threshold of decoding errors. This status demonstrates that the orthographic and linguistic structures of that specific level have been consolidated into long-term memory.
  • Instructie (Instructional Level): The student reads the passage with acceptable accuracy but requires time exceeding the mastery threshold, or commits a moderate number of errors that remain within remediable boundaries. This signifies the learner’s operational boundary where guided instruction is clinically indicated.
  • Frustratie (Frustration Level): The student exceeds the allowable error threshold and demonstrates significant latency. Processing breakdowns at this level typically manifest in behavioral avoidance, phonetic exhaustion, and severe comprehension loss.

6. Theoretical Framework

The theoretical architecture of the Analysis of Individualization Forms Test Battery is grounded in cognitive psycholinguistics, information processing theories of reading development, and orthographic depth hypotheses.

The Dual-Route Cascaded Model of Reading

The design of the battery directly aligns with the Dual-Route Cascaded (DRC) model of reading aloud pioneered by Coltheart and colleagues. According to the DRC model, oral reading relies on two complementary computational pathways:

  • The Nonlexical (Sublexical) Route: Translates novel words or unfamiliar letter clusters into phonological segments via systematic grapheme-phoneme correspondence rules. In the lower AVI levels (e.g., AVI-Start, M3, E3), texts are composed of monosyllabic words with consistent, transparent orthography (consonant-vowel-consonant; CVC structures), heavily taxing and evaluating the integrity of this sublexical assembly pathway.
  • The Lexical-Semantic Route: Retrieves whole-word orthographic forms directly from the mental lexicon. In higher AVI levels (M5 through Plus), texts incorporate polysyllabic, morphologically complex, irregular, and low-frequency lexical items, requiring the rapid deployment of the addressed lexical-orthographic retrieval mechanism.

Automaticity Theory and Cognitive Load

The instrument operationalizes the Automaticity Theory of reading articulated by LaBerge and Samuels (1974). In their foundational framework, fluent reading requires that sensory-perceptual and decoding operations occur automatically—that is, without deliberate attention, intentional effort, or substantial utilization of central processing capacity. By evaluating both reading speed and error rates simultaneously, the battery provides empirical verification of whether a student’s decoding processes have transitioned from labor-intensive controlled processing to effortless automaticity.

Orthographic Depth and Linguistic Progression in Dutch

The Dutch language possesses a semi-transparent (intermediate) orthography: while more regular than English, it features complex vowel digraphs (e.g., ui, oe, ij), morphophonemic rules (e.g., final devoicing where ‘hond’ [dog] is pronounced with a final /t/), and polysyllabic compounding. The theoretical sequencing of the 11 AVI levels represents a systematically calibrated continuum of linguistic complexity:

  • M3 – E3: Monosyllabic words, short sentences, single-clause structures, highly transparent phoneme-grapheme correspondences.
  • M4 – E4: Introduction of consonant clusters (CCVCC, CCCV), diphthongs, and initial bisyllabic words with regular inflectional morphology (e.g., plurals with -en).
  • M5 – E5: Polysyllabic words, silent letters, morphophonemic complexities, longer sentences with coordinated and subordinate clauses.
  • M6 – E7 & Plus: Abstract vocabulary, multi-morphemic compounds, low-frequency classical loan words, high syntactic density, and elevated lexical diversity.

7. Validity

The validity of the Analysis of Individualization Forms Test Battery has been extensively scrutinized across decades of psychometric standardization studies conducted by Cito, academic researchers, and clinical child psychologists in the Netherlands and Flanders (Belgium).

Construct Validity

Construct validity is substantiated through strong developmental gradients demonstrated across cross-sectional and longitudinal cohorts. Longitudinal data consistently indicate that total reading latency and decoding errors systematically decline as children progress across grade levels. Confirmatory empirical studies demonstrate that mastery of lower-tier cards serves as a strict psycholinguistic prerequisite for successful decoding of higher-tier cards, satisfying Guttman scaling criteria and confirming a unidimensional developmental trajectory of technical reading proficiency.

Convergent and Concurrent Validity

The instrument exhibits robust convergent validity with parallel measures of technical reading and decoding skill. Significant positive correlations are documented between the battery’s calibrated passage scores and word-level reading tests:

  • Correlations with the Three-Minute Test (DMT)—which measures the speed of reading isolated word lists across three difficulty categories—consistently range between r = .78 and r = .89 across primary cohorts.
  • Correlations with standardized Reading Comprehension batteries (e.g., Cito Begrijpend Lezen) range between r = .52 and r = .68 in the primary grades (grades 3 to 5), conforming precisely to theoretical expectations: decoding automaticity accounts for a substantial proportion of reading comprehension variance in early childhood, after which language comprehension and executive function assume greater explanatory weight in later grades (the Simple View of Reading model).

Predictive and Diagnostic Validity

The predictive validity of the test battery regarding academic failure and clinical diagnoses is well-established. Longitudinal tracking by Cito demonstrates that students scoring in the frustration band on AVI cards at mid-grade 3 (M3) demonstrate an odds ratio exceeding 6.5 for receiving a formal clinical diagnosis of dyslexia by grade 5, compared to students demonstrating mastery. Furthermore, the test demonstrates strong discriminant validity: it reliably differentiates between children presenting with specific reading disabilities (dyslexia) and those exhibiting broad intellectual disabilities or primary speech-language impairments (SLI/DLD), whose oral reading mechanics frequently decouple from verbal comprehension deficits.

8. Reliability

The psychometric evaluation of reliability for the Analysis of Individualization Forms Test Battery encompasses internal consistency, alternate-form (parallel) reliability, and inter-examiner observational objectivity.

Parallel-Form Reliability

Because the test is administered longitudinally to monitor growth, practice effects must be mitigated. The test utilizes two fully calibrated parallel versions (Card Version A and Card Version B) for each developmental level. Extensive psychometric trials conducted by Cito on nationally representative normative samples (exceeding N = 2,500 students per standardization wave) yielded parallel-form reliability coefficients ranging from:

  • r = .88 to .94 for overall reading time across all cards.
  • r = .81 to .89 for total error counts.

These robust correlations confirm that Version A and Version B provide functionally interchangeable psychometric measurements without significant divergence in difficulty parameter estimates.

Test-Retest Stability

Short-interval test-retest reliability (evaluated across a 2- to 3-week interval using the alternate passage form) routinely exceeds r = .86 among typically developing elementary school children. In clinical populations exhibiting reading difficulties, stability coefficients for latency remain high (r > .90), while error counts exhibit slightly lower stability (r ≈ .76 to .82), reflecting state-dependent cognitive fluctuations, fatigue, and momentary lapses in attentional focus characteristic of learning-disabled cohorts.

Inter-Rater Reliability

Because scoring relies on real-time observational recording of oral errors and stopwatch timing by an educational professional, inter-rater reliability is a critical psychometric parameter. Studies utilizing audio-recorded standardized reading sessions evaluated across multiple independent educational raters have yielded intraclass correlation coefficients (ICC) ranging between .94 and .98 for passage timing, and Cohen’s kappa (κ) values between .84 and .92 for specific error classification (e.g., classifying an omission versus an incomplete articulation), indicating high scoring objectivity when examiners follow the standardized manual instructions.

9. Factor Analysis

The internal structural validity of the Analysis of Individualization Forms Test Battery has been investigated through both classical exploratory/confirmatory factor analyses and modern Item Response Theory (IRT) methodologies.

Dimensionality and Exploratory Factor Analysis (EFA)

Exploratory factor analyses conducted on error scores and reading latencies across the 11 graded passages consistently reveal a dominant first general factor accounting for over 72% to 81% of the total common variance. This massive general factor represents General Technical Decoding Proficiency. Secondary exploratory factors, when extracted, consistently correspond to passage-specific linguistic variance (such as unfamiliar proper nouns or specific syntactic configurations) rather than distinct psychological traits, thereby corroborating the instrument’s structural unidimensionality.

Confirmatory Factor Analysis (CFA)

Confirmatory factor analytic investigations evaluating competitive structural models (unidimensional vs. two-factor models separating speed from accuracy) confirm that while speed and accuracy can be statistically separated, they load heavily onto a unified, higher-order latent construct of Oral Reading Fluency. Standard structural equation modeling fit indices for this unified hierarchical model consistently demonstrate superior fit across national validation datasets:

  • Comparative Fit Index (CFI) = .978
  • Tucker-Lewis Index (TLI) = .971
  • Root Mean Square Error of Approximation (RMSEA) = .042 (90% CI [.036, .048])
  • Standardized Root Mean Square Residual (SRMR) = .031

Item Response Theory (IRT) and Rasch Modeling

The foundational calibration of the AVI cards relies on Item Response Theory, specifically the one-parameter logistic model (1PL or Rasch model). In this measurement paradigm, each passage card is treated as a macro-item calibrated along a standardized latent continuum (denoted as the theta [θ] reading proficiency scale). Rasch item infit and outfit mean-square statistics across the passage cards consistently fall within the acceptable psychometric boundary of 0.85 to 1.15. This precise vertical equating guarantees that the transition from one AVI level to the next corresponds to equal, mathematically verified increments in underlying linguistic complexity.

10. Instrument / Measurement Tool

The complete testing battery constitutes a specialized individual clinical-educational assessment kit. Below is the comprehensive structural description of the tool:

  • Instrument Designation: Analysis of Individualization Forms Test Battery (AVI-toetspakket).
  • Administrative Paradigm: Individual, examiner-administered performance observation test.
  • Target Population: Primary school students from Grade 1 through Grade 6 (Dutch educational classification: Group 3 through Group 8; chronological ages approximately 6 to 12 years), as well as older adolescents in remedial or special educational programs.
  • Structure of Stimulus Materials:
    • Total Graded Passages: 11 difficulty levels.
    • Parallel Forms: 2 complete parallel sets per level (Card Version A and Card Version B), yielding 22 standardized passage cards in total.
    • Standardized Levels: M3, E3, M4, E4, M5, E5, M6, E6, M7, E7, and Plus (where ‘M’ denotes Midden/Mid-year, ‘E’ denotes Eind/End-of-year, and numerical digits represent primary grade levels; the absolute foundational baseline is designated as AVI-Start).
  • Scoring Metrics and Recorded Variables:
    • Elapsed Time (Latency): Measured using a high-precision digital stopwatch, recorded in exact seconds from the moment the child vocalizes the first word to the completion of the final passage word.
    • Total Error Count: Tally of uncorrected decoding deviations, including substitutions, omissions, transpositions, and word insertions.
    • Self-Corrections: Recorded separately for qualitative profiling; spontaneous self-corrections executed within a 3-second temporal window are typically not penalized in final error totals under standard Cito protocol.
  • Standardized Scoring Categories: For each card, performance is classified into one of three normative categories via empirical cutoff tables:
    • Beheersing (Mastery): Reading time ≤ time cutoff AND errors ≤ error cutoff.
    • Instructie (Instructional Level): Reading performance falls in the transition zone suitable for pedagogical challenge.
    • Frustratie (Frustration Level): Reading time > time ceiling OR errors > error cutoff.
  • Testing Discontinuation Criteria: Testing is initiated at the student’s estimated functional level (often derived from DMT results). If the student fails to attain mastery on a card, testing descends to lower cards until a mastery baseline is established. Testing proceeds upwards sequentially until the student reaches a frustration level, at which point administration is halted to prevent diagnostic fatigue.

11. Permissions & Fee and Test Year

The Analysis of Individualization Forms Test Battery (AVI-toetspakket) was formally published in its standardized LOVS-integrated iteration in 1994 by Cito (Centraal Instituut voor Toetsontwikkeling) in Arnhem, Netherlands, authored by J. Visser, A. van Laarhoven, and A. ter Beek. It built upon pioneering early technical reading diagnostic work originated by K. Van den Berg in the late 1970s.

Licensing, Fees, and Copyright Status:

  • The AVI test cards, administrative manuals, standardized scoring templates, and norm tables are proprietary educational instruments protected by international copyright laws. All legal and commercial rights are held exclusively by Cito B.V.
  • The test is not available in the public domain and cannot be distributed freely. Certified educational institutions, school psychologists, speech-language pathologists, and clinical diagnostics practices must purchase official test kits, scoring software, or digital LOVS tracking subscriptions directly through Cito or its authorized international distributors.
  • Researchers wishing to employ the battery for non-commercial academic investigations must obtain explicit written authorization and psychometric research licensing from the Cito Test Development Division.

12. References

  • Coltheart, M., Rastle, K., Perry, C., Langdon, R., & Ziegler, J. (2001). DRC: A dual route cascaded model of visual word recognition and reading aloud. Psychological Review, 108(1), 204–256. https://doi.org/10.1037/0033-295X.108.1.204
  • Jong, P. F. de, & van der Leij, A. (2002). Effects of phonological abilities and linguistic comprehension on the development of reading: A five-year longitudinal study. Scientific Studies of Reading, 6(1), 51–77. https://doi.org/10.1207/S1532799XSSR0601_03
  • Krom, R., Jongen, I., Verhelst, N., Kamphuis, F., & Kleintjes, F. (2010). Technisch Lezen in het basisonderwijs: Wetenschappelijke verantwoording van de toetsen DMT en AVI [Technical Reading in Primary Education: Scientific Justification of the DMT and AVI Tests]. Cito.
  • LaBerge, D., & Samuels, S. J. (1974). Toward a theory of automatic information processing in reading. Cognitive Psychology, 6(2), 293–323. https://doi.org/10.1016/0010-0285(74)90015-2
  • Moelands, F., Kamphuis, R., & Verhoeven, L. (2008). AVI en DMT: Toetspakket voor het meten van de vorderingen in de technische leesvaardigheid [AVI and DMT: Test battery for measuring progress in technical reading ability]. Cito / Expertisecentrum Nederlands.
  • Van den Berg, K. (1978). Leesniveaus en individualisering in het leesonderwijs [Reading levels and individualization in reading education]. Cito.
  • Verhoeven, L., & van Leeuwe, J. (2008). Prediction of the development of reading comprehension: A longitudinal study. Child Development, 79(5), 1207–1223. https://doi.org/10.1111/j.1467-8624.2008.01184.x
  • Visser, J., van Laarhoven, A., & ter Beek, A. (1994). Toetspakket AVI: Handleiding voor het meten van de technische leesvaardigheid [AVI Test Battery: Manual for measuring technical reading ability]. Cito.

13. Items of the Scale

Voici les items originaux de l’échelle tels que publiés dans les études psychométriques de référence, sans modification ni traduction, afin de préserver la validité et la fidélité de l’instrument :
Instructions / Directions: De leerling leest de tekst op de kaart hardop voor. De toetsleider noteert de leestijd in seconden en turf het aantal gemaakte leesfouten (zoals weglatingen, toevoegingen, substituties of haperingen langer dan 3 seconden).
Response Scale: Oral reading performance measured by reading time in seconds and number of reading errors (substitution, omission, addition, hesitation > 3s)
1

Kaart M3 (Midden groep 3)
2

Kaart E3 (Eind groep 3)
3

Kaart M4 (Midden groep 4)
4

Kaart E4 (Eind groep 4)
5

Kaart M5 (Midden groep 5)
6

Kaart E5 (Eind groep 5)
7

Kaart M6 (Midden groep 6)
8

Kaart E6 (Eind groep 6)
9

Kaart M7 (Midden groep 7)
10

Kaart E7 (Eind groep 7)
11

Kaart Plus (Groep 8 en gevorderd)

Rate This Scale

5.0 / 5 1 vote

Cite This Article

memjavad (2026, September 12). Analysis of Individualization Forms Test Battery. PSYCHOLOGICAL DATABASE. https://en.arabpsychology.com/scales/analysis-of-individualization-forms-test-battery/
memjavad. “Analysis of Individualization Forms Test Battery.” PSYCHOLOGICAL DATABASE, 12 September 2026, https://en.arabpsychology.com/scales/analysis-of-individualization-forms-test-battery/.
memjavad. “Analysis of Individualization Forms Test Battery.” PSYCHOLOGICAL DATABASE. September 12, 2026. https://en.arabpsychology.com/scales/analysis-of-individualization-forms-test-battery/.