We proposed a finite exponential-gamma frailty mixture (FEGFM) regression model for recurrent-event count data observed over a fixed follow-up window. The model targets settings with overdispersion, heavy upper tails, and latent subgroup structure that are not well-captured by a single Poisson or negative binomial regression. Each latent component has its own log-linear mean regression and gamma frailty parameter, while component membership depends on covariates through multinomial logistic gating. After integrating out frailty, each component follows a negative binomial distribution, so the overall model is a covariate-dependent finite mixture of negative binomial regressions. We derived the main distributional properties, identifiability conditions, and an expectation-maximization algorithm for maximum-likelihood estimation. Simulation results showed improved fit and tail calibration relative to Poisson and single-component negative binomial benchmarks, while also indicating weak identification for some parameters in finite samples. An application to physician-visit counts illustrates the practical value of separating within-class frailty from between-class latent heterogeneity.
Citation: Mohieddine Rahmouni. Frailty mixtures for count regression: separating within- and between-class heterogeneity[J]. AIMS Mathematics, 2026, 11(8): 23893-23920. doi: 10.3934/math.2026962
We proposed a finite exponential-gamma frailty mixture (FEGFM) regression model for recurrent-event count data observed over a fixed follow-up window. The model targets settings with overdispersion, heavy upper tails, and latent subgroup structure that are not well-captured by a single Poisson or negative binomial regression. Each latent component has its own log-linear mean regression and gamma frailty parameter, while component membership depends on covariates through multinomial logistic gating. After integrating out frailty, each component follows a negative binomial distribution, so the overall model is a covariate-dependent finite mixture of negative binomial regressions. We derived the main distributional properties, identifiability conditions, and an expectation-maximization algorithm for maximum-likelihood estimation. Simulation results showed improved fit and tail calibration relative to Poisson and single-component negative binomial benchmarks, while also indicating weak identification for some parameters in finite samples. An application to physician-visit counts illustrates the practical value of separating within-class frailty from between-class latent heterogeneity.
| [1] | A. C. Cameron, P. K. Trivedi, Regression analysis of count data, 2 Eds., Cambridge: Cambridge University Press, 2013. https://doi.org/10.1017/CBO9781139013567 |
| [2] | J. M. Hilbe, Negative binomial regression, 2 Eds., Cambridge: Cambridge University Press, 2011. https://doi.org/10.1017/CBO9780511973420 |
| [3] |
J. F. Lawless, Negative binomial and mixed poisson regression, Can. J. Stat., 15 (1987), 209–225. https://doi.org/10.2307/3314912 doi: 10.2307/3314912
|
| [4] | B. G. Lindsay, Mixture models: theory, geometry, and applications, Durham: Institute of Mathematical Statistics, 1995. https://doi.org/10.1214/cbms/1462106013 |
| [5] | G. J. McLachlan, D. Peel, Finite mixture models, New York: John Wiley & Sons, Inc., 2000. https://doi.org/10.1002/0471721182 |
| [6] |
P. Deb, P. Trivedi, Demand for medical care by the elderly: a finite mixture approach, J. Appl. Econ., 12 (1997), 313–336. https://doi.org/10.1002/(SICI)1099-1255(199705)12:3<313::AID-JAE440>3.0.CO;2-G doi: 10.1002/(SICI)1099-1255(199705)12:3<313::AID-JAE440>3.0.CO;2-G
|
| [7] |
P. Wang, M. L. Puterman, I. Cockburn, N. Le, Mixed Poisson regression models with covariate dependent rates, Biometrics, 52 (1996), 381–400. https://doi.org/10.2307/2532881 doi: 10.2307/2532881
|
| [8] |
B. Grün, F. Leisch, FlexMix version 2: finite mixtures with concomitant variables and varying and constant parameters, J. Stat. Softw., 28 (2008), 1–35. https://doi.org/10.18637/jss.v028.i04 doi: 10.18637/jss.v028.i04
|
| [9] |
D. Lambert, Zero-inflated Poisson regression, with an application to defects in manufacturing, Technometrics, 34 (1992), 1–14. https://doi.org/10.2307/1269547 doi: 10.2307/1269547
|
| [10] |
J. Mullahy, Specification and testing of some modified count data models, J. Econometrics, 33 (1986), 341–365. https://doi.org/10.1016/0304-4076(86)90002-3 doi: 10.1016/0304-4076(86)90002-3
|
| [11] |
Y. Wang, J. Li, Q. Zhou, W. Wang, A unified framework for complex survival data: accounting for clustering, cure fractions, and competing risks, BMC Med. Res. Methodol., 26 (2026), 105. https://doi.org/10.1186/s12874-026-02842-z doi: 10.1186/s12874-026-02842-z
|
| [12] |
F. Leisch, FlexMix: a general framework for finite mixture models and latent class regression in R, J. Stat. Softw., 11 (2004), 1–18. https://doi.org/10.18637/jss.v011.i08 doi: 10.18637/jss.v011.i08
|
| [13] | L. Duchateau, P. Janssen, The frailty model, New York: Springer, 2008. https://doi.org/10.1007/978-0-387-72835-3 |
| [14] | P. Hougaard, Analysis of multivariate survival data, New York: Springer, 2000. https://doi.org/10.1007/978-1-4612-1304-8 |
| [15] |
X. L. Meng, D. B. Rubin, Maximum likelihood estimation via the ECM algorithm: a general framework, Biometrika, 80 (1993), 267–278. https://doi.org/10.1093/biomet/80.2.267 doi: 10.1093/biomet/80.2.267
|
| [16] |
H. Chen, J. Chen, J. D. Kalbfleisch, A modified likelihood ratio test for homogeneity in finite mixture models, J. R. Stat. Soc. B, 63 (2001), 19–29. https://doi.org/10.1111/1467-9868.00273 doi: 10.1111/1467-9868.00273
|
| [17] |
H. Kasahara, K. Shimotsu, Non-parametric identification and estimation of the number of components in multivariate mixtures, J. R. Stat. Soc. B, 76 (2014), 97–111. https://doi.org/10.1111/rssb.12022 doi: 10.1111/rssb.12022
|
| [18] |
C. Hennig, Identifiability of models for clusterwise linear regression, J. Classif., 17 (2000), 273–296. https://doi.org/10.1007/s003570000022 doi: 10.1007/s003570000022
|
| [19] |
A. Dempster, N. Laird, D. Rubin, Maximum likelihood from incomplete data via the EM algorithm, J. R. Stat. Soc. B, 39 (1977), 1–22. https://doi.org/10.1111/j.2517-6161.1977.tb01600.x doi: 10.1111/j.2517-6161.1977.tb01600.x
|
| [20] |
B. G. Leroux, Consistent estimation of a mixing distribution, Ann. Stat., 20 (1992), 1350–1360. https://doi.org/10.1214/aos/1176348772 doi: 10.1214/aos/1176348772
|
| [21] |
J. Chen, P. Li, Hypothesis test for normal mixture models: the EM approach, Ann. Stat., 37 (2009), 2523–2542. https://doi.org/10.1214/08-AOS651 doi: 10.1214/08-AOS651
|
| [22] |
A. Xu, Y. Miu, J. Sun, S. Zhou, Y. Tang, A hierarchical Bayesian multivariate Wiener process model with dependent degradation rates and volatilities, IEEE Trans. Reliab., 75 (2026), 1420–1433. https://doi.org/10.1109/TR.2026.3674703 doi: 10.1109/TR.2026.3674703
|