Overdispersed and Zero-Truncated Count Models for Hospitality Insurance Claim Frequencies

Main Article Content

Wikanda Phaphan, Samach Sathitvudh, Peang-or Yeesa, Sudarat Nidsunkid

Abstract

Insurance claim frequencies are commonly modeled by Poisson regression, although the equidispersion assumption is often violated in real insurance portfolios. This paper studies count regression models for hospitality insurance claim counts under overdispersion and strictly positive sampling. We formulate the Poisson model as a benchmark and motivate the negative binomial model through its conditional mean-variance structure and Poisson–gamma mixture representation. Due to positive-count sampling, zero-truncated extensions are also contemplated. The empirical analysis is based on multi-year hospitality insurance claim data and statistically compares the benchmark, negative binomial along with Conway–Maxwell–Poisson, geometric, and zero-truncated specifications using dispersion diagnostics, likelihood-based criteria, and cross validation. The result shows that the negative binomial substantially improves the Poisson benchmark reducing the cross validated AIC by 1,003.46 on average. Among the extended models, the zero-truncated negative binomial model achieves the best likelihood-based fit, supported by the lowest AIC and BIC. These findings indicate that both unobserved heterogeneity and zero-truncation are crucial features of hospitality insurance claim frequencies.

Article Details

References

  1. E. Ohlsson, B. Johansson, Non-Life Insurance Pricing with Generalized Linear Models, Springer Berlin Heidelberg, 2010. https://doi.org/10.1007/978-3-642-10791-7.
  2. K. Antonio, E.A. Valdez, Statistical Concepts of a Priori and a Posteriori Risk Classification in Insurance, AStA Adv. Stat. Anal. 96 (2012), 187-224. https://doi.org/10.1007/s10182-011-0152-7.
  3. A.C. Cameron, P.K. Trivedi, Regression Analysis of Count Data, 2nd ed., Cambridge University Press, 2013. https://doi.org/10.1017/cbo9781139013567.
  4. J.A. Nelder, R.W.M. Wedderburn, Generalized Linear Models, J. R. Stat. Soc. Ser. A Gen. 135 (1972), 370-384. https://doi.org/10.2307/2344614.
  5. A.E. Renshaw, Modelling the Claims Process in the Presence of Covariates, ASTIN Bull. 24 (1994), 265-285. https://doi.org/10.2143/AST.24.2.2005070.
  6. S. Haberman, A.E. Renshaw, Generalized Linear Models and Actuarial Science, Statistician 45 (1996), 407-436. https://doi.org/10.2307/2988543.
  7. G. Dionne, C. Vanasse, Automobile Insurance Ratemaking in the Presence of Asymmetrical Information, J. Appl. Econom. 7 (1992), 149-165. https://doi.org/10.1002/jae.3950070204.
  8. M. Denuit, S. Lang, Non-Life Rate-Making with Bayesian GAMs, Insur. Math. Econ. 35 (2004), 627-647. https://doi.org/10.1016/j.insmatheco.2004.08.001.
  9. C. Gourieroux, J. Jasiak, Heterogeneous INAR(1) Model with Application to Car Insurance, Insur. Math. Econ. 34 (2004), 177-192. https://doi.org/10.1016/j.insmatheco.2003.11.005.
  10. F.A. Haight, Handbook of the Poisson Distribution, John Wiley and Sons, 1967.
  11. E.W. Frees, Regression Modeling with Actuarial and Financial Applications, Cambridge University Press, 2009. https://doi.org/10.1017/cbo9780511814372.
  12. D. Lord, F. Mannering, The Statistical Analysis of Crash-Frequency Data: A Review and Assessment of Methodological Alternatives, Transp. Res. Part A Policy Pract. 44 (2010), 291-305. https://doi.org/10.1016/j.tra.2010.02.001.
  13. J.M. Hilbe, Modeling Count Data, Cambridge University Press, 2014. https://doi.org/10.1017/cbo9781139236065.
  14. P. Shi, E.A. Valdez, Multivariate Negative Binomial Models for Insurance Claim Counts, Insur. Math. Econ. 55 (2014), 18-29. https://doi.org/10.1016/j.insmatheco.2013.11.011.
  15. A.L. Byers, H. Allore, T.M. Gill, P.N. Peduzzi, Application of Negative Binomial Modeling for Discrete Outcomes, J. Clin. Epidemiol. 56 (2003), 559-564. https://doi.org/10.1016/s0895-4356(03)00028-3.
  16. H. Liu, R.A. Davidson, D.V. Rosowsky, J.R. Stedinger, Negative Binomial Regression of Electric Power Outages in Hurricanes, J. Infrastruct. Syst. 11 (2005), 258-267. https://doi.org/10.1061/(ASCE)1076-0342(2005)11:4(258).
  17. J.-P. Boucher, M. Denuit, M. Guillén, Risk Classification for Claim Counts: A Comparative Analysis of Various Zero-Inflated Mixed Poisson and Hurdle Models, N. Am. Actuar. J. 11 (2007), 110-131. https://doi.org/10.1080/10920277.2007.10597487.
  18. M. Denuit, X. Maréchal, S. Pitrebois, J.-F. Walhin, Actuarial Modelling of Claim Counts: Risk Classification, Credibility and Bonus-Malus Systems, Wiley, 2007. https://doi.org/10.1002/9780470517420.
  19. S.C.K. Lee, Delta Boosting Implementation of Negative Binomial Regression in Actuarial Pricing, Risks 8 (2020), 19. https://doi.org/10.3390/risks8010019.
  20. A.F. Lukman, O. Albalawi, M. Arashi, J. Allohibi, A.A. Alharbi, R.A. Farghali, Robust Negative Binomial Regression via the Kibria-Lukman Strategy: Methodology and Application, Mathematics 12 (2024), 2929. https://doi.org/10.3390/math12182929.
  21. G. Tzougas, A. Pignatelli di Cerchiara, The Multivariate Mixed Negative Binomial Regression Model with an Application to Insurance a Posteriori Ratemaking, Insur. Math. Econ. 101 (2021), 602-625. https://doi.org/10.1016/j.insmatheco.2021.10.001.
  22. S. Sathitvudh, P. Wongwiwat, W. Phaphan, Enhancing Tourist Forecasting in Thailand's National Parks with Zero-Inflated Models, WSEAS Trans. Environ. Dev. 21 (2025), 284-292. https://doi.org/10.37394/232015.2025.21.25.
  23. J.F. Lawless, Negative Binomial and Mixed Poisson Regression, Can. J. Stat. 15 (1987), 209-225. https://doi.org/10.2307/3314912.
  24. S. Hu, T.B. Murphy, A. O'Hagan, Bivariate Gamma Mixture of Experts Models for Joint Insurance Claims Modeling, arXiv:1904.04699, 2019. https://doi.org/10.48550/arXiv.1904.04699.
  25. Z. Hossain, Maria, Analyzing Overdispersed Antenatal Care Count Data in Bangladesh: Mixed Poisson Regression with Individual-Level Random Effects, Austrian J. Stat. 50 (2021), 78-90. https://doi.org/10.17713/ajs.v50i4.1163.
  26. H. Akaike, A New Look at the Statistical Model Identification, IEEE Trans. Autom. Control 19 (1974), 716-723. https://doi.org/10.1109/TAC.1974.1100705.
  27. P. Stoica, Y. Selén, Model-Order Selection: A Review of Information Criterion Rules, IEEE Signal Process. Mag. 21 (2004), 36-47. https://doi.org/10.1109/MSP.2004.1311138.
  28. S.S. Wilks, The Large-Sample Distribution of the Likelihood Ratio for Testing Composite Hypotheses, Ann. Math. Stat. 9 (1938), 60-62. https://doi.org/10.1214/aoms/1177732360.
  29. J.T. Grogger, R.T. Carson, Models for Truncated Counts, J. Appl. Econom. 6 (1991), 225-238. https://doi.org/10.1002/jae.3950060302.
  30. K.F. Sellers, S. Borle, G. Shmueli, The COM-Poisson Model for Count Data: A Survey of Methods and Applications, Appl. Stoch. Model. Bus. Ind. 28 (2012), 104-116. https://doi.org/10.1002/asmb.918.
  31. G. Shmueli, T.P. Minka, J.B. Kadane, S. Borle, P. Boatwright, A Useful Distribution for Fitting Discrete Data: Revival of the Conway-Maxwell-Poisson Distribution, J. R. Stat. Soc. Ser. C Appl. Stat. 54 (2005), 127-142. https://doi.org/10.1111/j.1467-9876.2005.00474.x.
  32. K.F. Sellers, B. Premeaux, Conway-Maxwell-Poisson Regression Models for Dispersed Count Data, WIREs Comput. Stat. 13 (2021), e1533. https://doi.org/10.1002/wics.1533.