Research article

Inference for maximum-entropy models with moment uncertainty sets: Duality and stability

  • Published: 10 August 2026
  • MSC : 62F10, 62G20, 90C25, 94A17, 60E15

  • We study a robust formulation of the maximum-entropy principle in which moment information is specified through uncertainty sets that encode estimation error. The estimator selects the least-informative distribution consistent with a convex admissible set of moment vectors. We provide a self-contained duality theory for general compact convex uncertainty sets, including strong duality, existence and uniqueness of primal solutions, and an exponential-family characterization governed by a penalized log-partition dual objective. We derive quantitative stability bounds showing Lipschitz dependence of optimal dual parameters on the estimated moments and the uncertainty radius, and we translate these bounds into controls on Kullback–Leibler (KL) divergence, total variation distance, and errors of bounded forecasts. Data-driven uncertainty sets built from concentration inequalities yield finite-sample feasibility and oracle-type entropy guarantees. Applications to probabilistic forecasting and risk modeling are developed, together with explicit discrete examples illustrating entropy-driven shrinkage within uncertainty regions.

    Citation: Badr S. Alnssyan, Abdelaziz Alsubie, Sajad A. Sheikh, Javid Gani Dar. Inference for maximum-entropy models with moment uncertainty sets: Duality and stability[J]. AIMS Mathematics, 2026, 11(8): 24331-24352. doi: 10.3934/math.2026982

    Related Papers:

  • We study a robust formulation of the maximum-entropy principle in which moment information is specified through uncertainty sets that encode estimation error. The estimator selects the least-informative distribution consistent with a convex admissible set of moment vectors. We provide a self-contained duality theory for general compact convex uncertainty sets, including strong duality, existence and uniqueness of primal solutions, and an exponential-family characterization governed by a penalized log-partition dual objective. We derive quantitative stability bounds showing Lipschitz dependence of optimal dual parameters on the estimated moments and the uncertainty radius, and we translate these bounds into controls on Kullback–Leibler (KL) divergence, total variation distance, and errors of bounded forecasts. Data-driven uncertainty sets built from concentration inequalities yield finite-sample feasibility and oracle-type entropy guarantees. Applications to probabilistic forecasting and risk modeling are developed, together with explicit discrete examples illustrating entropy-driven shrinkage within uncertainty regions.



    加载中


    [1] C. E. Shannon, A mathematical theory of communication, Bell Syst. Tech. J., 27 (1948), 379–423,623–656. https://doi.org/10.1002/j.1538-7305.1948.tb00917.x doi: 10.1002/j.1538-7305.1948.tb00917.x
    [2] E. T. Jaynes, Information theory and statistical mechanics, Phys. Rev., 106 (1957), 620–630. https://doi.org/10.1103/PhysRev.106.620 doi: 10.1103/PhysRev.106.620
    [3] J. E. Shore, R. W. Johnson, Axiomatic derivation of the principle of maximum entropy and the principle of minimum cross-entropy, IEEE T. Inform. Theor., 26 (1980), 26–37. https://doi.org/10.1109/TIT.1980.1056144 doi: 10.1109/TIT.1980.1056144
    [4] I. Csiszár, Ⅰ-divergence geometry of probability distributions and minimization problems, Ann. Prob., 3 (1975), 146–158.
    [5] I. Csiszár, Information projections revisited, IEEE T. Inform. Theor., 49 (2003), 1474–1490. https://doi.org/10.1109/TIT.2003.810633 doi: 10.1109/TIT.2003.810633
    [6] J. M. Borwein, A. S. Lewis, Duality relationships for entropy-like minimization problems, SIAM J. Optim., 1 (1991), 191–205. https://doi.org/10.1137/0801014 doi: 10.1137/0801014
    [7] C. Léonard, Minimization of entropy functionals, J. Math. Anal. Appl., 346 (2008), 183–204. https://doi.org/10.1016/j.jmaa.2008.04.048 doi: 10.1016/j.jmaa.2008.04.048
    [8] T. Sutter, D. Sutter, P. M. Esfahani, J. Lygeros, Generalized maximum entropy estimation, J. Mach. Learn. Res., 20 (2019), 1–29.
    [9] A. Ben-Tal, L. El Ghaoui, A. Nemirovski, Robust Optimization, Princeton University Press, 2009.
    [10] D. Bertsimas, I. Popescu, Optimal inequalities in probability theory: A convex optimization approach, SIAM J. Optim., 15 (2005), 780–804. https://doi.org/10.1137/S1052623401399903 doi: 10.1137/S1052623401399903
    [11] E. Delage, Y. Ye, Distributionally robust optimization under moment uncertainty with application to data-driven problems, Oper. Res., 58 (2010), 595–612. https://doi.org/10.1287/opre.1090.0741 doi: 10.1287/opre.1090.0741
    [12] H. Rahimian, S. Mehrotra, Frameworks and results in distributionally robust optimization, Open J. Math. Optim., 3 (2022), 1–85. https://doi.org/10.5802/ojmo.15 doi: 10.5802/ojmo.15
    [13] D. Kuhn, S. Shafiee, W. Wiesemann, Distributionally robust optimization, Acta Numer., 34 (2025), 579–804. https://doi.org/10.1017/S0962492924000084 doi: 10.1017/S0962492924000084
    [14] P. M. Esfahani, D. Kuhn, Data-driven distributionally robust optimization using the Wasserstein metric: Performance guarantees and tractable reformulations, Math. Prog., 171 (2018), 115–166. https://doi.org/10.1007/s10107-017-1172-1 doi: 10.1007/s10107-017-1172-1
    [15] J. Blanchet, K. R. A. Murthy, Quantifying distributional model risk via optimal transport, Math. Oper. Res., 44 (2019), 565–600. https://doi.org/10.1287/moor.2018.0936 doi: 10.1287/moor.2018.0936
    [16] S. Shafieezadeh-Abadeh, D. Kuhn, P. M. Esfahani, Regularization via mass transportation, J. Mach. Learn. Res., 20 (2019), 1–68.
    [17] T. Gneiting, A. E. Raftery, Strictly proper scoring rules, prediction, and estimation, J. Am. Stat. Assoc., 102 (2007), 359–378. https://doi.org/10.1198/016214506000001437 doi: 10.1198/016214506000001437
    [18] R. T. Rockafellar, S. Uryasev, Optimization of conditional value-at-risk, J. Risk, 2 (2000), 21–41. https://doi.org/10.21314/JOR.2000.038 doi: 10.21314/JOR.2000.038
    [19] R. T. Rockafellar, Convex analysis, Princeton University Press, 1970.
    [20] C. L. Canonne, A short note on an inequality between KL and TV, arXiv preprint, 2022, arXiv: 2202.07198.
    [21] A. B. Owen, Empirical likelihood, Chapman & Hall/CRC, 2001.
    [22] W. Hoeffding, Probability inequalities for sums of bounded random variables, J. Am. Stat. Assoc., 58 (1963), 13–30.
    [23] S. Jiang, H. Xu, Q. Sun, S. Huang, Robust learning of minimum error entropy under heavy-tailed noise, J. Comput. Appl. Math., 487 (2026), 117686. https://doi.org/10.1016/j.cam.2026.117686 doi: 10.1016/j.cam.2026.117686
    [24] Q. Sun, Y. Zhou, S. Huang, Fast rates of exponential cost function, Stat. Papers, 66 (2025), 1–19. https://doi.org/10.1007/s00362-025-01672-3 doi: 10.1007/s00362-025-01672-3
  • Reader Comments
  • © 2026 the Author(s), licensee AIMS Press. This is an open access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0)
通讯作者: 陈斌, bchen63@163.com
  • 1. 

    沈阳化工大学材料科学与工程学院 沈阳 110142

  1. 本站搜索
  2. 百度学术搜索
  3. 万方数据库搜索
  4. CNKI搜索

Metrics

Article views(257) PDF downloads(16) Cited by(0)

Article outline

Figures and Tables

Figures(1)  /  Tables(1)

/

DownLoad:  Full-Size Img  PowerPoint
Return
Return

Catalog