Announcement

Collapse
No announcement yet.
X
  • Filter
  • Time
  • Show
Clear All
new posts

  • get AMEs by combining -margins- with some sort of parallel process?

    Dear Statalist:

    I am running a massive multinomial logistic regression using restricted data on an organization's servers. I'm trying to get average marginal effects.

    This is not my machine, I have to connect to it remotely and use their virtual machine. This means that upgrading to a Stata / MP that can use 16 (or 32, or 64) cores is not an option (unfortunately).

    I had an idea to use the user-written package -parallel- by Yon and Quistorff to try and speed it up. I do not know if this is possible, but if it is theoretically possible I would like to try.

    Some context:
    This multinomial logistic regression has 60 outcomes, and the sample I'm running it on has over six million respondents. The multinomial logistic regression usually finishes in about 12-15 hours, and if I get predicted values using -margins- with no standard errors at all (-nose- option) it usually takes about 45 hours. I am currently trying to get predicted values using margins with standard errors and it is definitely using every last drop of RAM to do it--currently about about 110 gigabytes of RAM. I suppose it might even take more if I had the RAM to spare. I'm also not sure when it will finish. Maybe tomorrow, maybe Monday, maybe never.

    However, something very different happened when I tried to get average marginal effects (AMEs) using -margins- and getting standard errors. The process started off by using up huge chunks of processing power (between 15% and 50% of the available CPU) and about 3 gigabytes of RAM for about a day. It then dropped down to about 5% to 20% of CPU, but increased to using about 6 gigabytes of RAM. And it kept doing this for about a week before I finally killed the process.

    I suppose I can just set it up to run and come back and check it every week until it's done, but I'm not sure it will ever finish that way. I'm also not sure why it's not using more available RAM. I'm hoping that maybe there's a way to parallelize the AMEs further, chopping up my RAM into 6 gigabyte chunks across multiple instances. But that depends very much on how -margins- calculates the AME, and it is possible I am misunderstanding that part.

    Has anyone tried doing something like this with -parallel-? Or anything else? If I'm patient enough eventually I can get predicted values without the SEs, but getting SEs for the AMEs seems to require something stronger.

    Thanks for your time,
    Jonathan

  • #2
    Do you have some minimal code example for us to work with? Combining parallel with margins has always been on my list, but in my view, this is not trivial. While you can easily split the sample and compute predictive values separately, how you handle the standard errors?
    What about using a random sub-sample as workaround for now? If it turns out that a 10% sample already has very low SEs, maybe you dont even need the full one.
    Last edited by Felix Bittmann; Yesterday, 12:34.
    Best wishes

    Stata 18.0 MP | ORCID | Google Scholar

    Comment


    • #3
      I can give an example of what the code looks like, maybe that would help? The data is restricted, but I can use a sample dataset and expand it so it is roughly the size of the dataset I am using.

      Please don't critique the "analysis", this is just to reproduce the size of the dataset and the sorts of things I'm asking margins to do for me.

      Code:
      clear
      sysuse nlsw88.dta
      
      mlogit occupation i.race##c.age i.race##c.grade
      
      est sto example_mlogit
      
      
      
      clear
      sysuse nlsw88.dta
      
      expand 1200 // This is just to simulate the size of the data set
      
      est restore example_mlogit
      
      estimates esample:
      
      margins, over(race) at(age=(35(9)44) grade=(12(4)16)) post
      I should note that this code eventually will resolve after a few days if I add the -nose- option. I'm looking to get the standard errors. I realize that with 2.5 million cases my standard errors will be nonexistent, but I think I need to give them anyway.


      I'm also planning on doing something this for average marginal effects: to get the AMEs for age and grade over race. That one doesn't seem to soak up nearly as much memory, though it's possible it just didn't get very far. I suppose it is possible that parallel processing it could speed that up too.

      Thanks,
      Jonathan




      Originally posted by Felix Bittmann View Post
      Do you have some minimal code example for us to work with? Combining parallel with margins has always been on my list, but in my view, this is not trivial. While you can easily split the sample and compute predictive values separately, how you handle the standard errors?
      What about using a random sub-sample as workaround for now? If it turns out that a 10% sample already has very low SEs, maybe you dont even need the full one.

      Comment

      Working...
      X