We study a robust formulation of the maximum-entropy principle in which moment information is specified through uncertainty sets that encode estimation error. The estimator selects the least-informative distribution consistent with a convex admissible set of moment vectors. We provide a self-contained duality theory for general compact convex uncertainty sets, including strong duality, existence and uniqueness of primal solutions, and an exponential-family characterization governed by a penalized log-partition dual objective. We derive quantitative stability bounds showing Lipschitz dependence of optimal dual parameters on the estimated moments and the uncertainty radius, and we translate these bounds into controls on Kullback–Leibler (KL) divergence, total variation distance, and errors of bounded forecasts. Data-driven uncertainty sets built from concentration inequalities yield finite-sample feasibility and oracle-type entropy guarantees. Applications to probabilistic forecasting and risk modeling are developed, together with explicit discrete examples illustrating entropy-driven shrinkage within uncertainty regions.
Citation: Badr S. Alnssyan, Abdelaziz Alsubie, Sajad A. Sheikh, Javid Gani Dar. Inference for maximum-entropy models with moment uncertainty sets: Duality and stability[J]. AIMS Mathematics, 2026, 11(8): 24331-24352. doi: 10.3934/math.2026982
We study a robust formulation of the maximum-entropy principle in which moment information is specified through uncertainty sets that encode estimation error. The estimator selects the least-informative distribution consistent with a convex admissible set of moment vectors. We provide a self-contained duality theory for general compact convex uncertainty sets, including strong duality, existence and uniqueness of primal solutions, and an exponential-family characterization governed by a penalized log-partition dual objective. We derive quantitative stability bounds showing Lipschitz dependence of optimal dual parameters on the estimated moments and the uncertainty radius, and we translate these bounds into controls on Kullback–Leibler (KL) divergence, total variation distance, and errors of bounded forecasts. Data-driven uncertainty sets built from concentration inequalities yield finite-sample feasibility and oracle-type entropy guarantees. Applications to probabilistic forecasting and risk modeling are developed, together with explicit discrete examples illustrating entropy-driven shrinkage within uncertainty regions.
| [1] |
C. E. Shannon, A mathematical theory of communication, Bell Syst. Tech. J., 27 (1948), 379–423,623–656. https://doi.org/10.1002/j.1538-7305.1948.tb00917.x doi: 10.1002/j.1538-7305.1948.tb00917.x
|
| [2] |
E. T. Jaynes, Information theory and statistical mechanics, Phys. Rev., 106 (1957), 620–630. https://doi.org/10.1103/PhysRev.106.620 doi: 10.1103/PhysRev.106.620
|
| [3] |
J. E. Shore, R. W. Johnson, Axiomatic derivation of the principle of maximum entropy and the principle of minimum cross-entropy, IEEE T. Inform. Theor., 26 (1980), 26–37. https://doi.org/10.1109/TIT.1980.1056144 doi: 10.1109/TIT.1980.1056144
|
| [4] | I. Csiszár, Ⅰ-divergence geometry of probability distributions and minimization problems, Ann. Prob., 3 (1975), 146–158. |
| [5] |
I. Csiszár, Information projections revisited, IEEE T. Inform. Theor., 49 (2003), 1474–1490. https://doi.org/10.1109/TIT.2003.810633 doi: 10.1109/TIT.2003.810633
|
| [6] |
J. M. Borwein, A. S. Lewis, Duality relationships for entropy-like minimization problems, SIAM J. Optim., 1 (1991), 191–205. https://doi.org/10.1137/0801014 doi: 10.1137/0801014
|
| [7] |
C. Léonard, Minimization of entropy functionals, J. Math. Anal. Appl., 346 (2008), 183–204. https://doi.org/10.1016/j.jmaa.2008.04.048 doi: 10.1016/j.jmaa.2008.04.048
|
| [8] | T. Sutter, D. Sutter, P. M. Esfahani, J. Lygeros, Generalized maximum entropy estimation, J. Mach. Learn. Res., 20 (2019), 1–29. |
| [9] | A. Ben-Tal, L. El Ghaoui, A. Nemirovski, Robust Optimization, Princeton University Press, 2009. |
| [10] |
D. Bertsimas, I. Popescu, Optimal inequalities in probability theory: A convex optimization approach, SIAM J. Optim., 15 (2005), 780–804. https://doi.org/10.1137/S1052623401399903 doi: 10.1137/S1052623401399903
|
| [11] |
E. Delage, Y. Ye, Distributionally robust optimization under moment uncertainty with application to data-driven problems, Oper. Res., 58 (2010), 595–612. https://doi.org/10.1287/opre.1090.0741 doi: 10.1287/opre.1090.0741
|
| [12] |
H. Rahimian, S. Mehrotra, Frameworks and results in distributionally robust optimization, Open J. Math. Optim., 3 (2022), 1–85. https://doi.org/10.5802/ojmo.15 doi: 10.5802/ojmo.15
|
| [13] |
D. Kuhn, S. Shafiee, W. Wiesemann, Distributionally robust optimization, Acta Numer., 34 (2025), 579–804. https://doi.org/10.1017/S0962492924000084 doi: 10.1017/S0962492924000084
|
| [14] |
P. M. Esfahani, D. Kuhn, Data-driven distributionally robust optimization using the Wasserstein metric: Performance guarantees and tractable reformulations, Math. Prog., 171 (2018), 115–166. https://doi.org/10.1007/s10107-017-1172-1 doi: 10.1007/s10107-017-1172-1
|
| [15] |
J. Blanchet, K. R. A. Murthy, Quantifying distributional model risk via optimal transport, Math. Oper. Res., 44 (2019), 565–600. https://doi.org/10.1287/moor.2018.0936 doi: 10.1287/moor.2018.0936
|
| [16] | S. Shafieezadeh-Abadeh, D. Kuhn, P. M. Esfahani, Regularization via mass transportation, J. Mach. Learn. Res., 20 (2019), 1–68. |
| [17] |
T. Gneiting, A. E. Raftery, Strictly proper scoring rules, prediction, and estimation, J. Am. Stat. Assoc., 102 (2007), 359–378. https://doi.org/10.1198/016214506000001437 doi: 10.1198/016214506000001437
|
| [18] |
R. T. Rockafellar, S. Uryasev, Optimization of conditional value-at-risk, J. Risk, 2 (2000), 21–41. https://doi.org/10.21314/JOR.2000.038 doi: 10.21314/JOR.2000.038
|
| [19] | R. T. Rockafellar, Convex analysis, Princeton University Press, 1970. |
| [20] | C. L. Canonne, A short note on an inequality between KL and TV, arXiv preprint, 2022, arXiv: 2202.07198. |
| [21] | A. B. Owen, Empirical likelihood, Chapman & Hall/CRC, 2001. |
| [22] | W. Hoeffding, Probability inequalities for sums of bounded random variables, J. Am. Stat. Assoc., 58 (1963), 13–30. |
| [23] |
S. Jiang, H. Xu, Q. Sun, S. Huang, Robust learning of minimum error entropy under heavy-tailed noise, J. Comput. Appl. Math., 487 (2026), 117686. https://doi.org/10.1016/j.cam.2026.117686 doi: 10.1016/j.cam.2026.117686
|
| [24] |
Q. Sun, Y. Zhou, S. Huang, Fast rates of exponential cost function, Stat. Papers, 66 (2025), 1–19. https://doi.org/10.1007/s00362-025-01672-3 doi: 10.1007/s00362-025-01672-3
|