Data ScienceQuantitative MethodsStatistics & Psychometrics

Akaike’s Information Criterion: Model Selection Guide

Akaike’s Information Criterion (AIC) is a foundational statistical metric that balances model fit and parsimony using information theory. Explore its mathematical foundations, calculations, and practical applications in model selection.

memjavad
PUBLISHED
Scientifically Reviewed · Dr. Marwa Abd-Alazim · October 6, 2026
Medically & Scientifically Reviewed Verified: October 6, 2026
Dr. Marwa Abd-Alazim Ph.D.
Professor of Psychology • University of Kerbala
Review Criteria & Clinical Standards

This content undergoes rigorous scientific peer-review and medical editorial standards at Arab Psychology Network to ensure clinical accuracy, validity, and compliance with evidence-based guidelines from leading psychological and healthcare authorities (APA / WHO).

Akaike’s Information Criterion stands as one of the most foundational and transformative principles in statistical inference, econometrics, psychometrics, and machine learning. By balancing goodness-of-fit against the mathematical cost of model complexity, it bridges theoretical information theory and applied statistical modeling. Understanding this criterion enables researchers to navigate the perpetual dilemma between model underfitting and excessive parameterization with mathematical rigor.

Akaike’s Information Criterion (AIC)

1. Concise Definition

Akaike’s Information Criterion (AIC) is an asymptotic, mathematical estimator of relative model quality that quantifies the trade-off between the goodness of fit of a statistical model and its parsimony. Formulated mathematically as AIC = 2k – 2ln(L), where k represents the number of estimated parameters and L signifies the maximized value of the likelihood function, the criterion assigns a numerical score to candidate models evaluated on a singular empirical dataset.

Rather than testing a formal null hypothesis in the classical Fisherian or Neyman-Pearson traditions, AIC provides a comparative, information-theoretic ranking mechanism. The absolute value of an AIC score holds no direct inferential meaning in isolation; rather, the relative differences between AIC values across distinct competing models serve as an estimate of the relative amount of information lost when a given model is employed to approximate the true, underlying, data-generating reality.

In practical research paradigms, the optimal model within an a priori designated candidate set is the one that minimizes the AIC score. By formalizing Occam’s razor into an objective mathematical metric, the criterion penalizes the inclusion of redundant parameters, thereby preventing researchers from overfitting sample noise while preserving genuine structural variance.

2. Etymology & Linguistic Origin

The term derives its patronymic designation from the Japanese statistician Hirotugu Akaike (1927–2009), who formulated the criterion in the early 1970s while working at the Institute of Statistical Mathematics in Tokyo. In his seminal 1973 and 1974 publications, Akaike initially titled the formulation simply “An Information Criterion,” which naturally yielded the acronym AIC. As the metric gained international renown, the statistical community universally retrofitted the acronym to honor its creator as “Akaike’s Information Criterion.”

Linguistically, the root word information stems from the Latin informare, meaning “to give form, shape, or represent mentally,” which later evolved via Old French into Middle English as a descriptor of communicated knowledge. In modern mathematical discourse, the concept was formalized through Claude Shannon’s information theory. The term criterion originates from the Ancient Greek kritērion (κριτήριον), denoting a standard, benchmark, or means by which something is judged or decided. Hence, Akaike’s Information Criterion translates literally and conceptually to an objective benchmark for adjudicating the informational integrity of mathematical constructs.

3. Pronunciation & Grammatical Form

In standard English academic discourse, the term is pronounced phonetically as /ah-kah-ee-keh-z ɪn-fər-meɪ-ʃən kraɪ-tɪər-i-ən/. The patronymic surname Akaike consists of four distinct syllables in Japanese (赤池, a-ka-i-ke), commonly pronounced with equal syllable weight.

Grammatically, the term functions as a singular proper noun phrase. It is universally abbreviated using the uppercase initialism AIC, which functions syntactically as a countable noun when referring to specific calculated scores (e.g., “the model demonstrated a lower AIC”) or as an uncountable mass noun when referencing the methodology itself (e.g., “model selection conducted via AIC”). Derived and related grammatical forms include the adjective AIC-based (e.g., “an AIC-based selection protocol”) and corrected forms such as AICc (small-sample corrected AIC) and QAIC (quasi-likelihood AIC).

4. Detailed Conceptual Explanation

At the center of scientific modeling lies an irreducible philosophical tension: true empirical processes in nature, behavior, biology, and society are vastly complex and effectively infinite-dimensional, whereas empirical models are deliberate, finite simplifications. When fitting a mathematical model to observed observations, introducing additional parameters almost invariably improves the empirical goodness of fit, raising the maximum likelihood or inflating the classical coefficient of determination (R²). However, this increase frequently captures sample-specific idiosyncratic variance—commonly termed noise—rather than underlying structural regularities.

This pathology is widely known as overfitting. A model afflicted with overfitting exhibits extraordinary fit within the observed training sample but fails catastrophically when tasked with predicting independent, unobserved observations. Conversely, an overly constrained model with excessively few parameters suffers from underfitting, failing to capture meaningful latent trends and systematic relationships. Akaike fundamentally reconceptualized model selection not as a quest to identify the “true” model, but as an optimization problem designed to minimize information loss when approximating reality.

The mathematical architecture of AIC is encapsulated in the canonical equation:

AIC = 2k – 2 ln(L)

Here, ln(L) represents the natural logarithm of the maximized likelihood function under the evaluated model. The term -2 ln(L) is an index of badness-of-fit: as the model fits the empirical data closer, the likelihood L increases, causing -2 ln(L) to decrease. Conversely, the term 2k acts as a strict, linear penalty scaled directly to the number of freely estimated parameters k. Every additional parameter estimated by the model must yield an improvement in log-likelihood of at least 1.0 unit simply to prevent the total AIC score from increasing. Consequently, AIC operationalizes a quantifiable, non-arbitrary tension between fidelity to the data and model simplicity.

Critically, AIC does not evaluate models on an absolute scale, nor does it generate an interpretable p-value. It is strictly a relative metric. When evaluating a pre-specified candidate set of models, researchers compute the AIC difference (ΔAIC) for each model relative to the candidate possessing the minimum score. By doing so, the framework provides an intuitive, continuous hierarchy of empirical plausibility, allowing scientists to assess competing theoretical representations without relying on nested model structures or conventional significance testing.

5. Historical Development

Prior to the early 1970s, the dominant paradigm for statistical model evaluation relied upon classical null hypothesis significance testing (NHST), forward and backward stepwise selection based on arbitrary alpha thresholds, and traditional residual variance measures. In 1951, Solomon Kullback and Richard Leibler published their landmark paper introducing what is now known as the Kullback-Leibler divergence (K-L information), a measure of the discrepancy between two probability distributions. For over two decades, however, K-L divergence remained an abstract theoretical concept because calculating it required complete knowledge of the true, unobservable reality.

Hirotugu Akaike achieved a profound breakthrough during the late 1960s and early 1970s while examining time series analysis and autoregressive modeling at the Institute of Statistical Mathematics. Akaike discovered that the maximized log-likelihood of a fitted model was an asymptotically biased estimator of the expected Kullback-Leibler information between the model and the true generating distribution. Most brilliantly, he proved that this asymptotic bias was approximately equal to the number of freely estimated parameters, k.

Akaike formally unveiled this discovery in 1973 at the Second International Symposium on Information Theory in Tsahkadsor, Armenia (then part of the Soviet Union), in a paper titled “Information Theory and an Extension of the Maximum Likelihood Principle.” He subsequently expanded the operational and theoretical architecture in his landmark 1974 paper, “A New Look at the Statistical Model Identification,” published in the IEEE Transactions on Automatic Control. This publication rapidly became one of the most cited works in mathematical statistics.

During the late 1980s and 1990s, statisticians like Clifford M. Hurvich and Chih-Ling Tsai (1989) identified that AIC suffered from substantial small-sample bias, frequently leading to overfitting when the ratio of observations to parameters was low. This work culminated in the development of the small-sample corrected criterion (AICc). Concurrently, ecologists and biostatisticians Kenneth P. Burnham and David R. Anderson published several influential texts, notably their 1998 and 2002 monographs, which popularized information-theoretic model selection and multimodel inference throughout the wider life and social sciences.

6. Theoretical Foundations

The axiomatic foundation of AIC rests upon the convergence of two major mathematical disciplines: Shannon-Weaver information theory and Fisherian maximum likelihood estimation. In information theory, the discrepancy between the true underlying probability distribution f(x) and an approximating model distribution g(x|θ) is defined by the Kullback-Leibler divergence, expressed continuously as:

I(f, g) = ∫ f(x) ln( f(x) / g(x|θ) ) dx

This integral can be algebraically decomposed into two discrete components:

I(f, g) = ∫ f(x) ln(f(x)) dx – ∫ f(x) ln(g(x|θ)) dx

The first term, ∫ f(x) ln(f(x)) dx, is the statistical entropy of the true reality. Because this reality is fixed and invariant across all candidate models, it represents an unknown constant, C. Consequently, minimizing the relative information loss between candidate models reduces entirely to maximizing the second term: the expected log-likelihood of the candidate model with respect to the true generating distribution, denoted as E_f[ln(g(X|θ))].

In empirical studies, researchers estimate the parameter vector θ by finding the maximum likelihood estimates θ̂ on a finite empirical sample. However, substituting θ̂ directly back into the log-likelihood function introduces a systemic positive bias, because the parameters have been specifically optimized to match the idiosyncratic peculiarities of that exact dataset. The model appears to perform better on its own data than it would on an independent replication drawn from the same distribution.

Akaike demonstrated via asymptotic Taylor series expansion that under regular conditions, the expected value of this optimism bias equals the dimension of the parameter vector, k. Multiplying by -2 for historical consistency with generalized deviance yields the foundational bias correction: -2E[ln(L)] ≈ -2ln(L̂) + 2k. Thus, minimizing AIC directly corresponds to selecting the model that minimizes the estimated expected information loss relative to the unknown truth.

7. Key Components, Types & Dimensions

The broader information-theoretic framework surrounding AIC encompasses several key components, variations, and derived metrics:

  • Goodness-of-Fit Component (-2 ln L): Quantifies how effectively the probability distribution specified by the parameter estimates reproduces the observed sample observations. Smaller values reflect superior empirical concordance.
  • Complexity Penalty (2k): Scales linearly with every freely estimated model parameter, including regression coefficients, variance components, and residual terms, acting as an explicit deterrent against unwarranted complexity.
  • Small-Sample Corrected AIC (AICc): Developed by Sugiura (1978) and Hurvich and Tsai (1989), AICc introduces an adjusted penalty term formulated as: AICc = AIC + [2k(k + 1)] / [n – k – 1], where n represents the total sample size. When n / k < 40, standard AIC tends to underestimate information loss, selecting models that are overly complex; AICc effectively eliminates this bias and asymptotically converges to AIC as sample sizes increase.
  • Akaike Differences (ΔAIC): Calculated as ΔAIC_i = AIC_i – AIC_min, this value reflects the relative performance gap between model i and the best-performing candidate. Guidelines developed by Burnham and Anderson establish that models with ΔAIC between 0 and 2 possess substantial empirical support; those between 4 and 7 exhibit considerably less support; and models with ΔAIC greater than 10 have essentially negligible support.
  • Akaike Weights (w_i): Normalized model probabilities calculated through the exponentiation of half the negative ΔAIC values: w_i = exp(-0.5 ΔAIC_i) / ∑ exp(-0.5 ΔAIC_j). These weights sum to 1.0 across the candidate set, representing the relative likelihood or strength of evidence that model i is the best approximating model within the candidate collection.
  • Quasi-AIC (QAIC): An adaptation tailored for overdispersed count or proportion data, introducing a variance inflation factor (ĉ) to prevent standard AIC from erroneously favoring overparameterized models due to unmodeled clustering or heterogeneity.

8. Examples & Illustrative Cases

To demonstrate the mechanics of AIC, consider a clinical psychologist investigating the determinants of treatment response in a major depressive disorder intervention trial involving 250 patients. The clinician hypothesizes several potential linear regression models to predict post-treatment symptom severity based on various baseline predictors, including baseline severity, cognitive reactivity scores, and sleep disturbance indices.

The clinician constructs four non-nested candidate models:

  • Model 1 (Intercept only): Estimates 1 mean parameter and 1 residual variance parameter (k = 2). Max log-likelihood = -520.4. AIC = 2(2) – 2(-520.4) = 1044.8.
  • Model 2 (Baseline Severity): Estimates 1 intercept, 1 regression slope, and 1 residual variance (k = 3). Max log-likelihood = -485.1. AIC = 2(3) – 2(-485.1) = 976.2.
  • Model 3 (Baseline Severity + Cognitive Reactivity): Estimates 1 intercept, 2 slopes, and 1 residual variance (k = 4). Max log-likelihood = -470.2. AIC = 2(4) – 2(-470.2) = 948.4.
  • Model 4 (Baseline Severity + Cognitive Reactivity + Sleep Disturbance + 5 Noise Variables): Estimates 1 intercept, 8 slopes, and 1 residual variance (k = 10). Max log-likelihood = -468.8. AIC = 2(10) – 2(-468.8) = 957.6.

Comparing the models shows that Model 4 achieves the highest raw log-likelihood (-468.8), indicating the closest absolute fit to the historical data. However, Model 4 required six additional parameters to gain only 1.4 units of log-likelihood over Model 3. Under AIC’s mathematical formulation, the parameter penalty (an increase of 12 points) outstripped the empirical gain (a decrease of 2.8 points). Consequently, Model 3 achieves an AIC of 948.4, outperforming Model 4 (AIC = 957.6) with a ΔAIC of 9.2. Model 3 is correctly selected as the optimal approximating model, successfully resisting the temptation to include redundant predictors that overfit sample noise.

9. Measurement & Assessment

Because AIC is an analytical mathematical construct, it is computed algorithmically via statistical software rather than measured through psychometric testing. Calculating AIC requires fitting a candidate model using numerical optimization procedures that maximize the log-likelihood function, such as the Newton-Raphson, Broyden-Fletcher-Goldfarb-Shanno (BFGS), or Expectation-Maximization (EM) algorithms.

Once the optimization routine converges upon the parameter vector θ̂, the exact maximized log-likelihood value is extracted. Concurrently, the degrees of freedom representing all estimated model parameters are compiled. In complex multi-level models, generalized additive models, or structural equation models, tracking k requires careful bookkeeping to ensure that auxiliary parameters—such as threshold boundaries in ordinal probit regressions, autoregressive error terms, and residual covariances—are accounted for alongside primary regression coefficients.

Researchers must also verify whether the software reports AIC using full log-likelihood expressions or formulations omitting constant terms. While omitting additive constants does not affect the calculation of ΔAIC within a single statistical package, cross-software comparisons can yield conflicting raw AIC values if one program retains theoretical integration constants while another discards them. Consistency across software platforms and candidate sets is therefore critical for reliable empirical evaluation.

10. Applications & Practical Significance

AIC is applied broadly across quantitative research disciplines. In psychometrics and structural equation modeling (SEM), researchers rely on AIC to adjudicate between competing latent factor structures. When evaluating whether a newly developed psychological assessment measures a unidimensional construct or a correlated multi-factor architecture, AIC allows investigators to compare non-nested models without the restrictive constraints of the standard chi-square difference test.

In clinical psychology and psychiatry, AIC guides the selection of growth curve trajectories in longitudinal trials. Investigators use it to determine whether symptom trajectories follow linear, quadratic, or piecewise spline trends across therapy sessions. Using AIC protects clinical research from overfitting idiosyncratic developmental patterns observed in small clinical samples.

In cognitive science and computational modeling, researchers deploy AIC to evaluate mathematical representations of cognitive processes, such as competing drift-diffusion models (DDMs) of decision-making or reinforcement learning algorithms. By applying AIC weights, cognitive scientists can systematically compare different cognitive architectures, evaluating how well each balances predictive capacity and computational parsimony.

Finally, in epidemiology and public health, AIC is central to the selection of time series and spatial lag models for disease transmission. It provides an objective basis for selecting lag lengths and climatic predictors in generalized linear models, helping prevent the overparameterization that can undermine forecasting models during health crises.

11. Research & Empirical Evidence

The empirical validity and operational boundaries of AIC have been thoroughly mapped through extensive Monte Carlo simulation studies and applied methodology over the past five decades. Research by Hurvich and Tsai (1989, 1993) demonstrated that in regression and autoregressive modeling where sample sizes are small relative to model parameters, standard AIC exhibits an empirical tendency toward selecting overparameterized models. Their work demonstrated that the corrected variant, AICc, virtually eliminates this finite-sample overfitting, outperforming uncorrected AIC across varying distributions and error structures.

A long-running debate in the literature examines the contrast between AIC and the Bayesian Information Criterion (BIC), developed by Gideon Schwarz in 1978. Extensive simulation studies by Kenneth Burnham, David Anderson, and Mark Berkman have shown that AIC and BIC optimize fundamentally different theoretical goals:

  • Efficiency vs. Consistency: AIC is asymptotically efficient; it minimizes mean squared error in prediction and does not assume that the true data-generating model exists within the candidate set.
  • Dimensionality Considerations: When the “true” model is assumed to be finite and simple, BIC consistently identifies that exact structure as the sample size approaches infinity. However, when the generating reality is infinite-dimensional and characterized by tapering effects (the reality typical of human behavior and biological systems), AIC consistently selects models that approximate the true process more effectively than BIC.

Empirical evaluations within structural equation modeling (e.g., Markland, 2007; Preacher & Merkle, 2012) further confirm that AIC differences are more reliable than traditional null-hypothesis tests when identifying the optimal number of latent factors, particularly in samples of moderate size where chi-square tests are overly sensitive to minor model mis-specifications.

12. Cultural & Cross-Cultural Considerations

While mathematical operations are culturally invariant, the interpretation, prioritization, and reporting of model selection criteria can reflect distinct scientific and institutional traditions. In Western psychological and behavioral sciences, researchers trained in classical Fisherian and Neyman-Pearson frameworks long favored hypothesis testing, treating the p-value as the gold standard for scientific validation. As a result, informational criteria like AIC were initially adopted slowly in disciplines where empirical rigor was historically equated with rejecting a null hypothesis.

In contrast, disciplines such as Japanese econometrics, Scandinavian quantitative biology, and modern computational ecology rapidly integrated information-theoretic frameworks, driven by an epistemological perspective that views all models as approximations rather than absolute truths. As cross-cultural behavioral research increasingly addresses complex, multi-site datasets, information criteria have gained widespread acceptance. They provide an objective framework for comparing structural models of psychological constructs across different cultural contexts without relying on sample-size-sensitive significance tests.

13. Criticisms, Debates & Limitations

Despite its mathematical foundations and widespread adoption, AIC is subject to several well-documented theoretical criticisms and practical limitations:

  • Lack of an Absolute Measure of Fit: AIC can only evaluate the relative quality of models within a specified candidate set. If every candidate model is poorly conceptualized, the model with the lowest AIC will merely be the best of an inadequate group. AIC does not provide an absolute metric of whether a model explains an acceptable proportion of observed variance.
  • Candidate Set Dependency: The validity of an AIC-based conclusion depends directly on the theoretical plausibility of the candidate set. P-hacking and post-hoc data dredging can occur if researchers construct dozens of exploratory combinations and selectively report the model with the lowest AIC.
  • Asymptotic Optimism in Small Samples: When the ratio of observations to parameters is low (e.g., n / k < 40), uncorrected AIC underestimates information loss and shows a documented bias toward overparameterized models, making the use of AICc essential in these settings.
  • Incomparability Across Differing Datasets: AIC cannot compare models estimated on different datasets or models evaluated on different subsets of observations. If missing observations vary across models due to listwise deletion, raw AIC comparisons become mathematically invalid.
  • Fixed-Penalty Critique: Unlike the Bayesian Information Criterion, whose penalty term incorporates sample size (ln(n)k), the AIC penalty remains fixed at 2k. Some statisticians criticize this property, pointing out that AIC is not asymptotically consistent: as the sample size approaches infinity, the probability that AIC selects an overparameterized model does not converge to zero.

14. Related Terms & Distinctions

A comprehensive understanding of AIC requires distinguishing it from related statistical metrics and model evaluation tools:

  • Bayesian Information Criterion (BIC): Formulated as BIC = k ln(n) – 2ln(L). While AIC approximates Kullback-Leibler divergence to maximize out-of-sample predictive accuracy, BIC approximates the integrated marginal likelihood from a Bayesian framework, assuming that the true data-generating model exists within the candidate set. Because ln(n) > 2 for any sample size n ≥ 8, BIC penalizes complexity more severely than AIC, routinely favoring simpler, more parsimonious models.
  • Deviance Information Criterion (DIC): A Bayesian generalization of AIC designed for models where parameters are estimated via Markov Chain Monte Carlo (MCMC) simulations, replacing k with an effective number of parameters (pD) derived from posterior distributions.
  • Widely Applicable Information Criterion (WAIC): Also known as the Watanabe-Akaike Information Criterion, WAIC extends information criteria into fully Bayesian frameworks by calculating out-of-sample expectation without requiring joint normality assumptions.
  • Likelihood Ratio Test (LRT): A classical hypothesis testing method used to compare two models. Unlike AIC, the Likelihood Ratio Test can only evaluate hierarchically nested models and relies on arbitrary alpha significance thresholds rather than a continuous information metric.
  • Cross-Validation (CV): A non-parametric resampling technique (such as leave-one-out cross-validation) that measures predictive accuracy by iteratively fitting and evaluating models on partitioned data subsets. Asymptotically, leave-one-out cross-validation is mathematically equivalent to model selection using AIC.

15. Summary / Key Takeaways

Akaike’s Information Criterion remains a cornerstone of modern statistical modeling. By combining Shannon’s information theory, Kullback-Leibler divergence, and maximum likelihood estimation, AIC transforms the philosophical principle of parsimony into an objective mathematical metric. Rather than testing whether an artificial null hypothesis can be rejected, AIC evaluates competing models by estimating how much information each loses when approximating an underlying reality.

The central take-home principles of AIC include:

  • AIC balances empirical goodness-of-fit against parameter complexity via the formula 2k – 2ln(L).
  • It is a strictly relative metric designed to compare candidate models evaluated on identical datasets; absolute AIC values hold no intrinsic meaning in isolation.
  • For small sample sizes (n / k < 40), the small-sample corrected metric (AICc) should be used to avoid bias toward overparameterized models.
  • Model comparisons should be reported using ΔAIC values and Akaike weights, providing a continuous assessment of empirical evidence across a candidate set.
  • AIC is mathematically equivalent to leave-one-out cross-validation, optimizing out-of-sample predictive performance rather than attempting to uncover a supposedly fixed, simple “true” model.

By moving statistical modeling away from binary significance thresholds and toward an information-theoretic assessment of competing explanations, Hirotugu Akaike provided scientists with a principled, mathematically elegant methodology for identifying the most effective empirical models of the real world.

References

  • Akaike, H. (1973). Information theory and an extension of the maximum likelihood principle. In B. N. Petrov & F. Csáki (Eds.), Second International Symposium on Information Theory (pp. 267–281). Akadémiai Kiadó.
  • Akaike, H. (1974). A new look at the statistical model identification. IEEE Transactions on Automatic Control, 19(6), 716–723. https://doi.org/10.1109/TAC.1974.1100705
  • Burnham, K. P., & Anderson, D. R. (2002). Model selection and multimodel inference: A practical information-theoretic approach (2nd ed.). Springer-Verlag. https://doi.org/10.1007/b97636
  • Hurvich, C. M., & Tsai, C.-L. (1989). Regression and time series model selection in small samples. Biometrika, 76(2), 297–307. https://doi.org/10.1093/biomet/76.2.297
  • Kullback, S., & Leibler, R. A. (1951). On information and sufficiency. The Annals of Mathematical Statistics, 22(1), 79–86. https://doi.org/10.1214/aoms/1177729694

Cite This Article

memjavad (2026, October 6). Akaike’s Information Criterion: Model Selection Guide. PSYCHOLOGICAL DATABASE. https://en.arabpsychology.com/dictionary/akaikes-information-criterion-aic/
memjavad. “Akaike’s Information Criterion: Model Selection Guide.” PSYCHOLOGICAL DATABASE, 6 October 2026, https://en.arabpsychology.com/dictionary/akaikes-information-criterion-aic/.
memjavad. “Akaike’s Information Criterion: Model Selection Guide.” PSYCHOLOGICAL DATABASE. October 6, 2026. https://en.arabpsychology.com/dictionary/akaikes-information-criterion-aic/.