We present a complete mathematical treatment of the reliability-weighted product-of-experts (RW-PoE) framework for multimodal Bayesian classification with provable guarantees. The modality posteriors are fused via a log-linear pool with weights $ {\mathit{\boldsymbol{w}}}\in\Delta_{M-1} $ fitted by maximum likelihood on the simplex. We prove the following: (i) Strict convexity of the weight objective if and only if the modality log-posteriors are independent in the simplex-tangent sense of Definition 4.2, which guarantees a unique global maximum likelihood estimate (MLE); (ii) Additive decomposition of $ H(Y)-H(Y\!\mid\! X) $ across a hierarchical cascade; (iii) An explicit non-asymptotic contraction bound on an interior simplex region and the exact $ \gamma $-dependent local linearization of the finite-step entropic mirror update; (iv) Finite-sample marginal conformal coverage at least $ 1-\alpha $, with the upper discretization bound $ 1-\alpha+1/(n_{\mathrm{cal}}+1) $ under the no-ties condition. A Hellinger–Bayes criterion yields a distribution-specific floor $ R^\star_{a, \mathrm{root}}\ge0.235 $ on the acoustic-only error at the root node for the observed distributions, showing that fusion is necessary in this setting (Remark 3.3 specifies the scope of this claim). A clinical study on $ n = 1105 $ infant cry episodes (CHU Sidi Bel-Abbès) reports $ \widehat{{\mathit{\boldsymbol{w}}}} = (0.42\pm0.03, \, 0.31\pm0.04, \, 0.27\pm0.03) $, realized coverage $ 0.910 $ (target $ 0.900 $; no-ties theoretical band $ [0.900, 0.905] $), accuracy $ 0.916 $, $ \kappa = 0.889 $ (95% CI: $ 0.844 $–$ 0.934 $).
Citation: Safia Benarbia, Fatimah Alshahrani, Mohammed Fethi Khalfi, Wahiba Bouabsa. A reliability-weighted product-of-experts framework for multimodal Bayesian classification: convexity, conformal coverage, and an information-theoretic necessity criterion[J]. AIMS Mathematics, 2026, 11(9): 31134-31184. doi: 10.3934/math.20261231
We present a complete mathematical treatment of the reliability-weighted product-of-experts (RW-PoE) framework for multimodal Bayesian classification with provable guarantees. The modality posteriors are fused via a log-linear pool with weights $ {\mathit{\boldsymbol{w}}}\in\Delta_{M-1} $ fitted by maximum likelihood on the simplex. We prove the following: (i) Strict convexity of the weight objective if and only if the modality log-posteriors are independent in the simplex-tangent sense of Definition 4.2, which guarantees a unique global maximum likelihood estimate (MLE); (ii) Additive decomposition of $ H(Y)-H(Y\!\mid\! X) $ across a hierarchical cascade; (iii) An explicit non-asymptotic contraction bound on an interior simplex region and the exact $ \gamma $-dependent local linearization of the finite-step entropic mirror update; (iv) Finite-sample marginal conformal coverage at least $ 1-\alpha $, with the upper discretization bound $ 1-\alpha+1/(n_{\mathrm{cal}}+1) $ under the no-ties condition. A Hellinger–Bayes criterion yields a distribution-specific floor $ R^\star_{a, \mathrm{root}}\ge0.235 $ on the acoustic-only error at the root node for the observed distributions, showing that fusion is necessary in this setting (Remark 3.3 specifies the scope of this claim). A clinical study on $ n = 1105 $ infant cry episodes (CHU Sidi Bel-Abbès) reports $ \widehat{{\mathit{\boldsymbol{w}}}} = (0.42\pm0.03, \, 0.31\pm0.04, \, 0.27\pm0.03) $, realized coverage $ 0.910 $ (target $ 0.900 $; no-ties theoretical band $ [0.900, 0.905] $), accuracy $ 0.916 $, $ \kappa = 0.889 $ (95% CI: $ 0.844 $–$ 0.934 $).
| [1] |
C. Genest, J. V. Zidek, Combining probability distributions: a critique and an annotated bibliography, Stat. Sci., 1 (1986), 114–135. https://doi.org/10.1214/ss/1177013825 doi: 10.1214/ss/1177013825
|
| [2] |
A. E. Abbas, A Kullback–Leibler view of linear and log-linear pools, Decis. Anal., 6 (2009), 25–37. https://doi.org/10.1287/deca.1080.0133 doi: 10.1287/deca.1080.0133
|
| [3] |
A. P. Dempster, N. M. Laird, D. B. Rubin, Maximum likelihood from incomplete data via the EM algorithm, J. R. Stat. Soc. Ser. B, 39 (1977), 1–22. https://doi.org/10.1111/j.2517-6161.1977.tb01600.x doi: 10.1111/j.2517-6161.1977.tb01600.x
|
| [4] |
T. A. Louis, Finding the observed information matrix when using the EM algorithm, J. R. Stat. Soc. Ser. B, 44 (1982), 226–233. https://doi.org/10.1111/j.2517-6161.1982.tb01203.x doi: 10.1111/j.2517-6161.1982.tb01203.x
|
| [5] | L. Devroye, L. Györfi, G. Lugosi, A probabilistic theory of pattern recognition, Springer, 1996. https://doi.org/10.1007/978-1-4612-0711-5 |
| [6] | V. Vovk, A. Gammerman, G. Shafer, Algorithmic learning in a random world, Springer, 2005. https://doi.org/10.1007/b106715 |
| [7] |
A. N. Angelopoulos, S. Bates, Conformal prediction: a gentle introduction, Found. Trends Mach. Learn., 16 (2023), 494–591. https://doi.org/10.1561/2200000101 doi: 10.1561/2200000101
|
| [8] | T. Heskes, Selecting weighting factors in logarithmic opinion pools, Proceedings of the 1997 Conference on Advances in Neural Information Processing Systems, 1998, 266–272. https://doi.org/10.5555/302528.302604 |
| [9] |
G. E. Hinton, Training products of experts by minimizing contrastive divergence, Neural Comput., 14 (2002), 1771–1800. https://doi.org/10.1162/089976602760128018 doi: 10.1162/089976602760128018
|
| [10] |
R. A. Jacobs, M. I. Jordan, S. J. Nowlan, G. E. Hinton, Adaptive mixtures of local experts, Neural Comput., 3 (1991), 79–87. https://doi.org/10.1162/neco.1991.3.1.79 doi: 10.1162/neco.1991.3.1.79
|
| [11] | Y. Romano, M. Sesia, E. J. Candès, Classification with valid and adaptive coverage, arXiv, 2020. https://doi.org/10.48550/arXiv.2006.02544 |
| [12] | A. N. Angelopoulos, S. Bates, M. I. Jordan, J. Malik, Uncertainty sets for image classifiers using conformal prediction, Int. Conf. Learn. Representations, 2021. |
| [13] | T. M. Cover, J. A. Thomas, Elements of information theory, 2 Eds., John Wiley & Sons, Inc., 2006. https://doi.org/10.1002/047174882X |
| [14] |
S. Watanabe, Information theoretical analysis of multivariate correlation, IBM J. Res. Dev., 4 (1960), 66–82. https://doi.org/10.1147/rd.41.0066 doi: 10.1147/rd.41.0066
|
| [15] | A. Bhattacharyya, On a measure of divergence between two statistical populations defined by their probability distributions, Bull. Calcutta Math. Soc., 35 (1943), 99–109. |
| [16] |
E. R. DeLong, D. M. DeLong, D. L. Clarke-Pearson, Comparing the areas under two or more correlated receiver operating characteristic curves: a nonparametric approach, Biometrics, 44 (1988), 837–845. https://doi.org/10.2307/2531595 doi: 10.2307/2531595
|
| [17] |
H. Lu, R. M. Freund, Y. Nesterov, Relatively smooth convex optimization by first-order methods, and applications, SIAM J. Optim., 28 (2018), 333–354. https://doi.org/10.1137/16M1099546 doi: 10.1137/16M1099546
|