Research article Special Issues

Information criteria as statistical complexity measures for AI models

  • Published: 11 August 2026
  • MSC : 62F07; 62J02; 62F15; 62B10

  • Artificial intelligence (AI) models are often described as complex because they contain many parameters, use nonlinear architectures, or require substantial computational resources. Architectural size, however, is not the same as statistical complexity. A large model can generalize well when regularization or optimization restricts the functions it effectively uses, while a smaller model may overfit if its selected structure is unstable. This paper develops a focused framework for interpreting information criteria as measures of effective statistical complexity in AI model evaluation. It connects Akaike's information criterion, the Bayesian information criterion, Takeuchi's information criterion, the minimum description length principle, and the widely applicable information criterion to distinct sources of model complexity: explicit parameters, shrinkage, structural partitioning, and singular or distributed representations. The framework is positioned relative to distribution-free capacity measures such as VC dimension and Rademacher complexity, and its applicability conditions are stated explicitly. The paper argues against a single universal complexity score and instead recommends reporting three separable quantities: goodness of fit, complexity cost, and residual uncertainty. A numerical illustration on a simulated regression problem shows how the three-part decomposition separates architectural size from generalization behavior across linear, penalized, and tree-based models. This decomposition provides a practical basis for comparing models, assessing parsimony, and avoiding unsupported claims based only on raw parameter counts or in-sample fit.

    Citation: Mohieddine Rahmouni. Information criteria as statistical complexity measures for AI models[J]. AIMS Mathematics, 2026, 11(8): 24650-24666. doi: 10.3934/math.2026992

    Related Papers:

  • Artificial intelligence (AI) models are often described as complex because they contain many parameters, use nonlinear architectures, or require substantial computational resources. Architectural size, however, is not the same as statistical complexity. A large model can generalize well when regularization or optimization restricts the functions it effectively uses, while a smaller model may overfit if its selected structure is unstable. This paper develops a focused framework for interpreting information criteria as measures of effective statistical complexity in AI model evaluation. It connects Akaike's information criterion, the Bayesian information criterion, Takeuchi's information criterion, the minimum description length principle, and the widely applicable information criterion to distinct sources of model complexity: explicit parameters, shrinkage, structural partitioning, and singular or distributed representations. The framework is positioned relative to distribution-free capacity measures such as VC dimension and Rademacher complexity, and its applicability conditions are stated explicitly. The paper argues against a single universal complexity score and instead recommends reporting three separable quantities: goodness of fit, complexity cost, and residual uncertainty. A numerical illustration on a simulated regression problem shows how the three-part decomposition separates architectural size from generalization behavior across linear, penalized, and tree-based models. This decomposition provides a practical basis for comparing models, assessing parsimony, and avoiding unsupported claims based only on raw parameter counts or in-sample fit.



    加载中


    [1] H. Akaike, A new look at the statistical model identification, IEEE Trans. Automat. Contr., 19 (1974), 716–723. https://doi.org/10.1109/TAC.1974.1100705 doi: 10.1109/TAC.1974.1100705
    [2] G. Schwarz, Estimating the dimension of a model, Ann. Stat., 6 (1978), 461–464. https://doi.org/10.1214/aos/1176344136 doi: 10.1214/aos/1176344136
    [3] S. Hochreiter, J. Schmidhuber, Flat minima, Neural Comput., 9 (1997), 1–42. https://doi.org/10.1162/neco.1997.9.1.1 doi: 10.1162/neco.1997.9.1.1
    [4] P. Bartlett, A. Montanari, A. Rakhlin, Deep learning: a statistical viewpoint, Acta Numer., 30 (2021), 87–201. https://doi.org/10.1017/S0962492921000027 doi: 10.1017/S0962492921000027
    [5] L. Breiman, Random forests, Mach. Learn., 45 (2001), 5–32. https://doi.org/10.1023/A:1010933404324 doi: 10.1023/A:1010933404324
    [6] B. Efron, The estimation of prediction error: covariance penalties and cross-validation, J. Am. Stat. Assoc., 99 (2004), 619–632. https://doi.org/10.1198/016214504000000692 doi: 10.1198/016214504000000692
    [7] H. Zou, T. Hastie, R. Tibshirani, On the "degrees of freedom" of the Lasso, Ann. Stat., 35 (2007), 2173–2192. https://doi.org/10.1214/009053607000000127 doi: 10.1214/009053607000000127
    [8] V. N. Vapnik, Statistical learning theory, New York: Wiley, 1998.
    [9] P. Bartlett, S. Mendelson, Rademacher and Gaussian complexities: risk bounds and structural results, J. Mach. Learn. Res., 3 (2002), 463–482.
    [10] M. Belkin, D. Hsu, S. Ma, S. Mandal, Reconciling modern machine-learning practice and the classical bias-variance trade-off, Proc. Nat. Acad. Sci. U.S.A., 116 (2019), 15849–15854. https://doi.org/10.1073/pnas.1903070116 doi: 10.1073/pnas.1903070116
    [11] O. Bousquet, A. Elisseeff, Stability and generalization, J. Mach. Learn. Res., 2 (2002), 499–526.
    [12] Y. Jiang, B. Neyshabur, H. Mobahi, D. Krishnan, S. Bengio, Fantastic generalization measures and where to find them, Proceedings of International Conference on Learning Representations 2020, 2020, 1–33.
    [13] C. Hurvich, C. Tsai, Regression and time series model selection in small samples, Biometrika, 76 (1989), 297–307. https://doi.org/10.1093/biomet/76.2.297 doi: 10.1093/biomet/76.2.297
    [14] J. Rissanen, Modeling by shortest data description, Automatica, 14 (1978), 465–471. https://doi.org/10.1016/0005-1098(78)90005-5 doi: 10.1016/0005-1098(78)90005-5
    [15] P. Grünwald, The minimum description length principle, Cambridge: MIT Press, 2007. https://doi.org/10.7551/mitpress/4643.001.0001
    [16] S. Watanabe, Asymptotic equivalence of Bayes cross validation and widely applicable information criterion in singular learning theory, J. Mach. Learn. Res., 11 (2010), 3571–3594.
    [17] A. Hoerl, R. Kennard, Ridge regression: biased estimation for nonorthogonal problems, Technometrics, 12 (1970), 55–67. https://doi.org/10.1080/00401706.1970.10488634 doi: 10.1080/00401706.1970.10488634
    [18] R. Tibshirani, Regression shrinkage and selection via the Lasso, J. R. Stat. Soc. B, 58 (1996), 267–288. https://doi.org/10.1111/j.2517-6161.1996.tb02080.x doi: 10.1111/j.2517-6161.1996.tb02080.x
    [19] H. Zou, T. Hastie, Regularization and variable selection via the elastic net, J. R. Stat. Soc. B, 67 (2005), 301–320. https://doi.org/10.1111/j.1467-9868.2005.00503.x doi: 10.1111/j.1467-9868.2005.00503.x
    [20] J. Friedman, Greedy function approximation: a gradient boosting machine, Ann. Stat., 29 (2001), 1189–1232. https://doi.org/10.1214/aos/1013203451 doi: 10.1214/aos/1013203451
    [21] T. Chen, C. Guestrin, XGBoost: a scalable tree boosting system, Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2016,785–794. https://doi.org/10.1145/2939672.2939785 doi: 10.1145/2939672.2939785
    [22] S. Watanabe, Algebraic geometry and statistical learning theory, Cambridge: Cambridge University Press, 2009. https://doi.org/10.1017/CBO9780511800474
    [23] D. McAllester, Some PAC-Bayesian theorems, Mach. Learn., 37 (1999), 355–363. https://doi.org/10.1023/A:1007618624809 doi: 10.1023/A:1007618624809
    [24] C. Zhang, S. Bengio, M. Hardt, B. Recht, O. Vinyals, Understanding deep learning requires rethinking generalization, Proceedings of International Conference on Learning Representations 2017, 2017, 1–15.
  • Reader Comments
  • © 2026 the Author(s), licensee AIMS Press. This is an open access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0)
通讯作者: 陈斌, bchen63@163.com
  • 1. 

    沈阳化工大学材料科学与工程学院 沈阳 110142

  1. 本站搜索
  2. 百度学术搜索
  3. 万方数据库搜索
  4. CNKI搜索

Metrics

Article views(263) PDF downloads(34) Cited by(0)

Article outline

Figures and Tables

Tables(3)

Other Articles By Authors

/

DownLoad:  Full-Size Img  PowerPoint
Return
Return

Catalog