Distributional Frailty Mixtures for Overdispersed Count Regression

Main Article Content

Mohieddine Rahmouni

Abstract

Standard finite mixtures of negative binomial regressions allow covariate-dependent means within latent classes but impose constant dispersion within each component. We propose the distributional finite exponential–gamma frailty mixture (D-FEGFM) regression model, which allows component-specific dispersion to depend on covariates through a log-linear regression. The model nests the constant-dispersion mixture and single-component distributional negative binomial regression. We derive the likelihood, establish identifiability up to label switching, extend the within/between variance decomposition to covariate-varying dispersion, and develop a generalized EM algorithm. A Monte Carlo study shows that the model is identifiable but weakly powered when mean and dispersion share the same covariates. In healthcare utilisation data, the strongest gains arise when demographic variables drive dispersion and clinical variables drive the mean, improving fit and held-out predictive performance, and yielding an interpretable insurance effect on utilisation regularity.

Article Details

References

  1. J.F. Lawless, Negative Binomial and Mixed Poisson Regression, Can. J. Stat. 15 (1987), 209–225. https://doi.org/10.2307/3314912.
  2. A.C. Cameron, P.K. Trivedi, Regression Analysis of Count Data, Cambridge University Press, 2013. https://doi.org/10.1017/cbo9781139013567.
  3. J.M. Hilbe, Negative Binomial Regression, Cambridge University Press, 2011. https://doi.org/10.1017/CBO9780511973420.
  4. G. McLachlan, D. Peel, Finite Mixture Models, Wiley, 2000. https://doi.org/10.1002/0471721182.
  5. P. Deb, P.K. Trivedi, Demand for Medical Care by the Elderly: A Finite Mixture Approach, J. Appl. Econom. 12 (1997), 313–336. https://doi.org/10.1002/(SICI)1099-1255(199705)12:3<313::AID-JAE440>3.0.CO;2-G.
  6. P. Wang, M.L. Puterman, I. Cockburn, N. Le, Mixed Poisson Regression Models with Covariate Dependent Rates, Biometrics 52 (1996), 381–400. https://doi.org/10.2307/2532881.
  7. B. Grün, F. Leisch, FlexMix Version 2: Finite Mixtures with Concomitant Variables and Varying and Constant Parameters, J. Stat. Softw. 28 (2008), 1–35. https://doi.org/10.18637/jss.v028.i04.
  8. G.K. Smyth, Generalized Linear Models with Varying Dispersion, J. R. Stat. Soc. Ser. B Stat. Methodol. 51 (1989), 47–60. https://doi.org/10.1111/j.2517-6161.1989.tb01747.x.
  9. R.A. Rigby, D.M. Stasinopoulos, Generalized Additive Models for Location, Scale and Shape, J. R. Stat. Soc. Ser. C Appl. Stat. 54 (2005), 507–554. https://doi.org/10.1111/j.1467-9876.2005.00510.x.
  10. M.D. Stasinopoulos, R.A. Rigby, G.Z. Heller, V. Voudouris, F. De Bastiani, Flexible Regression and Smoothing: Using GAMLSS in R, Chapman and Hall/CRC, 2017. https://doi.org/10.1201/b21973.
  11. R.A. Rigby, M.D. Stasinopoulos, G.Z. Heller, F. De Bastiani, Distributions for Modeling Location, Scale, and Shape: Using GAMLSS in R, Chapman and Hall/CRC, 2019. https://doi.org/10.1201/9780429298547.
  12. A. Herschtal, Poisson Beta Regression for Count Data with an Application to Hospital Length of Stay Data, Stat. Med. 44 (2025), e70217. https://doi.org/10.1002/sim.70217.
  13. C.F.J. Wu, On the Convergence Properties of the EM Algorithm, Ann. Stat. 11 (1983), 95–103. https://doi.org/10.1214/aos/1176346060.
  14. M. Rahmouni, Frailty Mixtures for Count Regression: Separating Within- and Between-Class Heterogeneity, AIMS Math. 11 (2026), 23893–23920. https://doi.org/10.3934/math.2026962.
  15. A.P. Dempster, N.M. Laird, D.B. Rubin, Maximum Likelihood from Incomplete Data via the EM Algorithm, J. R. Stat. Soc. Ser. B Stat. Methodol. 39 (1977), 1–22. https://doi.org/10.1111/j.2517-6161.1977.tb01600.x.
  16. R.M. Neal, G.E. Hinton, A View of the EM Algorithm that Justifies Incremental, Sparse, and Other Variants, in: M.I. Jordan (Ed.), Learning in Graphical Models, Springer Netherlands, Dordrecht, pp. 355–368, (1998). https://doi.org/10.1007/978-94-011-5014-9_12.
  17. H. Teicher, Identifiability of Finite Mixtures, Ann. Math. Stat. 34 (1963), 1265–1269. https://doi.org/10.1214/aoms/1177703862.
  18. S.J. Yakowitz, J.D. Spragins, On the Identifiability of Finite Mixtures, Ann. Math. Stat. 39 (1968), 209–214. https://doi.org/10.1214/aoms/1177698520.
  19. C. Hennig, Identifiability of Models for Clusterwise Linear Regression, J. Classif. 17 (2000), 273–296. https://doi.org/10.1007/s003570000022.
  20. H. Chen, J. Chen, J.D. Kalbfleisch, A Modified Likelihood Ratio Test for Homogeneity in Finite Mixture Models, J. R. Stat. Soc. Ser. B Stat. Methodol. 63 (2001), 19–29. https://doi.org/10.1111/1467-9868.00273.
  21. T.A. Louis, Finding the Observed Information Matrix When Using the EM Algorithm, J. R. Stat. Soc. Ser. B Stat. Methodol. 44 (1982), 226–233. https://doi.org/10.1111/j.2517-6161.1982.tb01203.x.