Announcement

Collapse
No announcement yet.
X
  • Filter
  • Time
  • Show
Clear All
new posts

  • Advice on what model to use


    I am trying to model the effect of incarceration on total out-of-pocket dental spending. When I take the log of total out-of-pocket dental spending and generate a histogram, the distribution shows a large number of zeros. These zeros likely arise from two distinct processes: individuals who had zero out-of-pocket expenditures because they did not visit the dentist and individuals who had zero out-of-pocket expenditures because they had dental insurance.

    To address this, I applied a selection CQR model using the user-generated command arhomme, but I am unsure if this is the best approach for handling the excess zeros. Would a Zero-Inflated Poisson (ZIP) model or another zero-inflated model be more appropriate? I have also included a list of variables I am considering for this analysis.


    Code:
    * Example generated by -dataex-. For more info, type help dataex
    clear
    input float(log_avrg_cost inc_d endentulism race age_cat) byte(male education veteran) float mothered byte wealth float(smoke_now chronicdisease) byte r11dentst
     5.860786 0 0 1 3 0 5 0 1 4 0 3 1
            0 0 0 1 1 0 5 0 1 4 0 3 1
     6.478509 0 0 1 2 0 3 0 0 4 0 0 1
     7.824446 0 0 1 2 0 4 0 1 4 1 0 1
     6.398595 0 0 1 3 1 5 1 1 4 0 2 1
     8.699764 0 0 1 2 0 5 0 0 4 0 0 1
     5.525453 0 0 4 3 0 5 0 1 4 0 2 1
            0 0 0 3 3 0 1 0 . 2 0 2 0
     6.685861 0 0 3 3 0 1 0 0 1 0 2 1
     6.398595 0 0 1 2 0 1 0 1 4 0 1 1
     7.378384 0 0 1 4 1 5 1 0 4 0 2 1
            0 0 0 3 3 0 1 0 0 1 0 2 1
            0 0 1 4 3 1 1 0 0 1 0 3 0
            0 0 1 2 3 0 3 0 0 1 0 1 1
            0 0 0 2 3 1 1 0 . 1 0 0 0
     7.151485 0 0 1 3 0 4 0 1 1 0 1 1
            0 1 1 2 3 0 1 0 . 1 0 1 0
     6.216606 0 0 1 3 0 4 0 1 4 0 0 1
     5.993961 0 0 1 3 1 5 1 1 4 0 0 1
      8.03948 0 0 1 3 0 1 0 1 4 0 1 1
            0 0 0 1 3 0 3 0 1 1 0 2 0
     5.303305 0 1 4 3 0 5 0 0 4 0 2 1
     5.303305 0 0 4 3 1 5 1 0 4 0 3 0
     4.620059 0 0 1 3 0 3 0 0 1 0 0 1
            0 0 1 1 3 1 4 1 . 2 0 1 0
     7.313887 0 0 1 2 0 5 0 1 2 0 0 1
     8.881975 0 0 1 4 0 5 0 1 4 0 2 1
     5.993961 0 0 1 3 1 5 1 1 4 0 2 1
     7.824446 0 0 1 3 0 3 0 1 4 0 0 1
     7.523481 0 0 1 3 1 5 0 0 4 0 3 1
     6.803505 0 0 1 3 0 5 0 1 4 0 1 1
            0 0 0 2 3 0 1 0 0 3 0 2 1
            0 0 0 2 3 0 5 0 1 1 0 2 0
            0 0 1 2 3 0 1 0 0 4 0 1 0
            0 0 1 2 2 0 3 0 0 1 0 0 0
     6.398595 0 0 2 3 1 3 1 1 3 0 3 1
            0 0 0 2 3 0 3 0 1 3 0 0 1
     6.908755 0 0 2 3 1 3 0 0 2 . 2 0
     5.786897 0 0 3 1 0 3 0 1 4 0 0 1
     4.836282 0 0 4 3 0 1 0 0 4 0 0 1
     5.303305 0 0 4 4 1 5 0 0 4 0 0 1
     6.908755 0 0 3 3 1 4 1 1 4 0 1 1
     3.314186 0 0 1 2 0 4 0 1 4 0 0 1
      7.31422 0 0 1 2 0 4 0 0 4 0 1 1
     8.987322 0 0 2 3 0 5 0 0 1 0 4 1
            0 0 1 4 3 0 4 0 0 2 0 0 0
     4.912655 1 0 1 3 1 1 0 1 2 0 1 1
      5.86221 0 0 1 3 0 5 0 1 2 0 1 1
     5.239098 0 0 1 3 1 3 1 0 4 0 2 1
            0 0 0 2 3 0 4 0 0 3 0 3 0
     6.803505 0 0 1 4 1 5 1 1 4 0 2 1
     7.438972 0 0 3 3 0 3 0 0 3 0 1 1
            0 0 1 3 3 1 2 1 0 3 0 1 0
    4.6151204 0 0 4 4 1 5 0 0 2 0 1 0
    3.9318256 0 0 4 3 0 5 0 0 2 0 2 1
            0 0 0 2 3 0 1 0 0 1 1 2 0
     5.351858 0 0 2 3 0 3 0 0 1 0 1 1
            0 0 0 2 3 0 1 0 . 1 0 2 0
      6.53814 0 0 1 3 1 4 1 0 1 0 3 1
     7.313887 1 0 4 2 0 1 0 1 2 0 6 1
     8.101981 0 0 1 4 1 3 1 1 4 0 4 1
      6.29803 0 0 1 3 0 1 0 1 4 0 1 1
     6.055613 0 0 1 3 0 5 0 1 3 0 1 1
            0 0 0 1 3 0 3 0 1 2 0 2 0
      7.91972 0 0 1 4 0 4 0 0 4 0 3 1
     6.908755 0 0 1 4 1 5 1 . 4 0 1 1
     7.003974 0 0 1 4 1 4 1 1 4 0 1 1
     6.763885 0 0 1 3 0 3 0 0 4 0 4 1
            0 0 0 1 2 0 3 0 1 2 0 1 1
            0 0 0 1 3 0 2 0 0 3 0 0 1
    1.7917595 0 0 3 3 0 3 0 . 1 0 3 1
            0 1 0 2 1 0 1 0 0 1 1 1 0
            0 0 0 2 1 0 4 0 0 1 1 2 0
            0 0 1 2 3 0 4 0 1 2 0 3 0
            0 0 0 1 3 0 1 0 0 1 0 2 0
     6.685861 0 0 1 3 1 4 1 1 2 0 1 1
     5.463832 0 0 1 3 1 3 0 . 3 0 2 1
            0 0 0 1 3 0 3 0 1 4 0 1 0
     4.620059 0 0 1 4 1 1 1 0 3 0 0 1
     6.634634 0 0 1 3 1 5 1 1 4 0 2 1
     5.525453 0 0 1 3 1 5 0 0 4 0 2 0
     5.860786 0 0 1 1 0 5 0 1 4 0 0 1
            0 0 1 1 3 1 2 0 0 3 0 3 0
     6.246107 0 0 3 2 0 1 0 0 3 . 2 1
            0 0 0 1 4 1 1 0 0 2 0 3 0
     5.673323 0 0 1 4 0 4 0 0 3 0 0 1
            0 0 1 2 3 0 3 0 0 1 0 3 0
            0 0 0 2 3 1 2 0 1 1 1 0 0
            0 1 1 2 3 1 3 0 1 2 1 0 0
            0 0 0 2 3 0 4 0 0 2 0 2 0
     3.433987 0 0 2 3 1 4 1 0 3 0 2 1
    3.9318256 1 0 2 1 0 4 0 0 3 1 2 1
     5.170484 0 0 2 3 1 1 0 1 2 0 3 1
            0 0 0 2 2 1 3 1 0 1 . 4 1
            0 0 1 1 3 1 1 1 0 3 0 4 0
     7.601402 0 0 1 3 0 1 0 0 4 0 0 1
            0 0 1 2 3 1 1 1 . 2 0 3 0
            0 0 0 2 2 0 3 0 1 3 0 1 0
            0 0 1 1 2 0 3 0 . 2 1 2 0
            0 1 1 1 3 1 4 0 1 4 0 4 0
    end
    label values inc_d inc_d
    label def inc_d 0 "No", modify
    label def inc_d 1 "Yes", modify
    label values endentulism endentulism
    label def endentulism 0 "No", modify
    label def endentulism 1 "Yes", modify
    label values race race
    label def race 1 "White", modify
    label def race 2 "Black", modify
    label def race 3 "Hispanic", modify
    label def race 4 "Other", modify
    label values age_cat age_cat
    label def age_cat 1 "50-59", modify
    label def age_cat 2 "60-69", modify
    label def age_cat 3 "70-79", modify
    label def age_cat 4 "80+", modify
    label values male male
    label def male 0 "Female", modify
    label def male 1 "Male", modify
    label values education EDUC
    label def EDUC 1 "1.lt high-school", modify
    label def EDUC 2 "2.ged", modify
    label def EDUC 3 "3.high-school graduate", modify
    label def EDUC 4 "4.some college", modify
    label def EDUC 5 "5.college and above", modify
    label values veteran veteran
    label def veteran 0 "No", modify
    label def veteran 1 "Yes", modify
    label values mothered mothered
    label def mothered 0 "Less than High School", modify
    label def mothered 1 "High School or Higher", modify
    label values smoke_now smoke_now
    label def smoke_now 0 "Non-Smoker", modify
    label def smoke_now 1 "Currently Smokes", modify
    label values r11dentst YESNO
    label def YESNO 0 "0.no", modify
    label def YESNO 1 "1.yes", modify


    Click image for larger version

Name:	image.png
Views:	1
Size:	69.9 KB
ID:	1772168

    Last edited by Luis Mijares Castaneda; 05 Feb 2025, 10:08.

  • #2
    I recommend you read Health Econometrics Using Stata (Deb, Norton and Manning, Stata Press, 2017). In Chapter 7 which is titled 'Models for continuous outcomes with mass at zero', the authors discussed a lot on the topic that is relevant to your question. And they demonstrated how to handle these (skewed continuous outcomes with many zeros) data using two-part models and generalized tobit selection model (heckman model). Let me cite a part of their introduction in Chapter 7.

    ... some important research questions in health economics and health policy involve only expenditures for those who spend at least some money, many more research questions involve health expenditures that include a substantial fraction of zeros—with the remaining values being positive, continuous, and severely skewed. For example, annual hospital expenditures are zero for most people but positive and often large for the subset who require hospital care. The majority of adults are nonsmokers, with many moderate smokers and a few heavy smokers. For any measure of healthcare use—inpatient, outpatient, emergency room, dental visit, preventive care—there is always a sizable fraction of the general population who do not use any healthcare during a defined period. The domains of all of these healthcare outcomes are either zero or positive. Statistical models that reflect the point mass at zero may better describe the relationships between the explanatory variables and the outcomes. (Deb, norton and Manning, 2017)
    https://www.stata.com/bookstore/heal...s-using-stata/

    Comment


    • #3
      Luis:
      as an aside to Chen's helpful reply, see also, in the same textbook, paragraph 8.4.2 on Zero-inflated models and the John Mullahy 's paramount contribution on this topic (HETEROGENEITY, EXCESS ZEROS, AND THE STRUCTURE OF COUNT DATA MODELS).
      Kind regards,
      Carlo
      (Stata 19.0)

      Comment

      Working...
      X