Dear Statalist:
I am running a massive multinomial logistic regression using restricted data on an organization's servers. I'm trying to get average marginal effects.
This is not my machine, I have to connect to it remotely and use their virtual machine. This means that upgrading to a Stata / MP that can use 16 (or 32, or 64) cores is not an option (unfortunately).
I had an idea to use the user-written package -parallel- by Yon and Quistorff to try and speed it up. I do not know if this is possible, but if it is theoretically possible I would like to try.
Some context:
This multinomial logistic regression has 60 outcomes, and the sample I'm running it on has over six million respondents. The multinomial logistic regression usually finishes in about 12-15 hours, and if I get predicted values using -margins- with no standard errors at all (-nose- option) it usually takes about 45 hours. I am currently trying to get predicted values using margins with standard errors and it is definitely using every last drop of RAM to do it--currently about about 110 gigabytes of RAM. I suppose it might even take more if I had the RAM to spare. I'm also not sure when it will finish. Maybe tomorrow, maybe Monday, maybe never.
However, something very different happened when I tried to get average marginal effects (AMEs) using -margins- and getting standard errors. The process started off by using up huge chunks of processing power (between 15% and 50% of the available CPU) and about 3 gigabytes of RAM for about a day. It then dropped down to about 5% to 20% of CPU, but increased to using about 6 gigabytes of RAM. And it kept doing this for about a week before I finally killed the process.
I suppose I can just set it up to run and come back and check it every week until it's done, but I'm not sure it will ever finish that way. I'm also not sure why it's not using more available RAM. I'm hoping that maybe there's a way to parallelize the AMEs further, chopping up my RAM into 6 gigabyte chunks across multiple instances. But that depends very much on how -margins- calculates the AME, and it is possible I am misunderstanding that part.
Has anyone tried doing something like this with -parallel-? Or anything else? If I'm patient enough eventually I can get predicted values without the SEs, but getting SEs for the AMEs seems to require something stronger.
Thanks for your time,
Jonathan
I am running a massive multinomial logistic regression using restricted data on an organization's servers. I'm trying to get average marginal effects.
This is not my machine, I have to connect to it remotely and use their virtual machine. This means that upgrading to a Stata / MP that can use 16 (or 32, or 64) cores is not an option (unfortunately).
I had an idea to use the user-written package -parallel- by Yon and Quistorff to try and speed it up. I do not know if this is possible, but if it is theoretically possible I would like to try.
Some context:
This multinomial logistic regression has 60 outcomes, and the sample I'm running it on has over six million respondents. The multinomial logistic regression usually finishes in about 12-15 hours, and if I get predicted values using -margins- with no standard errors at all (-nose- option) it usually takes about 45 hours. I am currently trying to get predicted values using margins with standard errors and it is definitely using every last drop of RAM to do it--currently about about 110 gigabytes of RAM. I suppose it might even take more if I had the RAM to spare. I'm also not sure when it will finish. Maybe tomorrow, maybe Monday, maybe never.
However, something very different happened when I tried to get average marginal effects (AMEs) using -margins- and getting standard errors. The process started off by using up huge chunks of processing power (between 15% and 50% of the available CPU) and about 3 gigabytes of RAM for about a day. It then dropped down to about 5% to 20% of CPU, but increased to using about 6 gigabytes of RAM. And it kept doing this for about a week before I finally killed the process.
I suppose I can just set it up to run and come back and check it every week until it's done, but I'm not sure it will ever finish that way. I'm also not sure why it's not using more available RAM. I'm hoping that maybe there's a way to parallelize the AMEs further, chopping up my RAM into 6 gigabyte chunks across multiple instances. But that depends very much on how -margins- calculates the AME, and it is possible I am misunderstanding that part.
Has anyone tried doing something like this with -parallel-? Or anything else? If I'm patient enough eventually I can get predicted values without the SEs, but getting SEs for the AMEs seems to require something stronger.
Thanks for your time,
Jonathan

Comment