<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/">
	<channel>
		<title>Statalist - Forums for Discussing Stata</title>
		<link>https://www.statalist.org/forums/</link>
		<description>Talk about all things related to Stata</description>
		<language>en</language>
		<lastBuildDate>Thu, 08 Oct 2026 07:44:48 GMT</lastBuildDate>
		<generator>vBulletin</generator>
		<ttl>60</ttl>
		<image>
			<url>images/misc/rss.png</url>
			<title>Statalist - Forums for Discussing Stata</title>
			<link>https://www.statalist.org/forums/</link>
		</image>
		<item>
			<title>New on SSC: rcsplot — publication-ready restricted cubic spline plots with one command</title>
			<link>https://www.statalist.org/forums/forum/general-stata-discussion/general/1787451-new-on-ssc-rcsplot-—-publication-ready-restricted-cubic-spline-plots-with-one-command</link>
			<pubDate>Wed, 07 Oct 2026 19:22:43 GMT</pubDate>
			<description>I am pleased to announce that rcsplot is now available from the SSC archive. To install: 
 
ssc install rcsplot 
rcsplot fits a restricted cubic...</description>
			<content:encoded><![CDATA[I am pleased to announce that rcsplot is now available from the SSC archive. To install:<br />
<br />
ssc install rcsplot<br />
rcsplot fits a restricted cubic spline (RCS) model for a continuous exposure and produces a publication-ready dose–response curve in a single command. It works with logistic, logit, poisson, stcox, and regress, and returns odds ratios, hazard ratios, rate ratios, or beta coefficients depending on the model.<br />
<br />
Key features:<ul><li><b>Reference-centred curves</b> — the curve is anchored at a user-specified reference (median by default), so the effect estimate is exactly 1 (or 0 for linear models) at that point.</li>
<li><b>Two P-values reported</b> — overall association and non-linearity, printed in the output and optionally on the plot.</li>
<li><b>Built-in trimming</b> — observations outside the 1st–99th percentile of the exposure are dropped before fitting, and the plotted x-range matches the modelled sample by default.</li>
<li><b>Survey support</b> — works with svy, including domain estimation via svyopts(, subpop(if ...)), with knots and reference computed on the trimmed subpopulation.</li>
<li><b>Publication-ready graphics</b> — optional histogram overlay, shaded confidence bands in Nature-journal style (ciband), and adjustable band opacity (cibopacity()).</li>
</ul>A typical call:<br />
<br />
webuse nhanes2, clear<br />
<br />
rcsplot logistic highbp zinc age i.sex bmi, knots(4) ciband histogram <br />
<br />
This fits a 4-knot RCS logistic model for zinc (adjusted for age, sex, and BMI), plots the OR curve with a shaded 95% CI, overlays a histogram of zinc.<br />
<br />
The command is documented with a full help file (help rcsplot).<br />
<br />
Feedback and bug reports are welcome.]]></content:encoded>
			<category domain="https://www.statalist.org/forums/forum/general-stata-discussion/general">General</category>
			<dc:creator>Zumin Shi</dc:creator>
			<guid isPermaLink="true">https://www.statalist.org/forums/forum/general-stata-discussion/general/1787451-new-on-ssc-rcsplot-—-publication-ready-restricted-cubic-spline-plots-with-one-command</guid>
		</item>
		<item>
			<title>Prediction Models and Stabilised Inverse-Probability of Treatment Weighting</title>
			<link>https://www.statalist.org/forums/forum/general-stata-discussion/general/1787446-prediction-models-and-stabilised-inverse-probability-of-treatment-weighting</link>
			<pubDate>Wed, 07 Oct 2026 05:55:09 GMT</pubDate>
			<description>Hi Statalist, 
 
I am trying to build a prediction model to estimate the true burden of respiratory syncytial virus (RSV), the leading cause of...</description>
			<content:encoded><![CDATA[Hi Statalist,<br />
<br />
I am trying to build a prediction model to estimate the true burden of respiratory syncytial virus (RSV), the leading cause of hospitalisation in young children, in Western Australia (WA).<br />
<br />
I have access to linked WA hospitalisation data, including perinatal, emergency department, vaccination, mortality, and laboratory pathology datasets. For each hospitalisation, we define a laboratory-confirmed RSV-associated admission as a hospital admission occurring within 4 days of the specimen date for a positive RSV test.<br />
<br />
A challenge is that RSV testing is relatively uncommon. Approximately 10% of the hospitalised children are tested for RSV, and among those tested, around 20% are RSV-positive.<br />
<br />
Variables:<ul><li>Predictors of interest include:<ul><li>Maternal characteristics ($mat_vars): E.g., age, smoking during pregnancy, plurality, etc.</li>
<li>Child characteristics ($child_vars): E.g., sex, gestational age at birth, birthweight, etc.</li>
<li>Clinical outcomes ($clin_vars): E.g., bronchiolitis, pneumonia, etc.</li>
<li>Time-varying predictors ($time_vars): E.g., month/season/year of birth, month/season/year of admission, etc.</li>
<li>Geographical predictors ($geo_vars): E.g., socioeconomic status, remoteness</li>
</ul></li>
<li>Additional variables:<ul><li>rsvtest (yes/no): Whether the child was tested for RSV</li>
<li>rsvpos (yes/no): RSV positivity among tested children</li>
</ul></li>
</ul>Proposed general steps of prediction model (in Stata):<ol class="decimal"><li>Fit an elastic-net logistic model among RSV-tested children to select predictors. I am planning to use multiple different alphas, however, for simplicity, I will use an alpha of 1 (i.e., LASSO).<ul><li>
<div class="bbcode_container">
	<div class="bbcode_description">Code:</div>
	<pre class="bbcode_code">elasticnet logit rsvpos $mat_vars $child_vars $clin_vars $time_vars $geo_vars if rsvtest==1, rseed(123) alphas(1) selection(cv) cluster(child_ID) nolog</pre>
</div></li>
</ul></li>
<li>Document the variables selected by elastic-net logistic model:<ul><li>
<div class="bbcode_container">
	<div class="bbcode_description">Code:</div>
	<pre class="bbcode_code">global EN_vars_a1 &quot;EN_vars&quot;</pre>
</div></li>
</ul></li>
<li>Use the selected variables in a model for the predicted probability of RSV testing and derive stabilised inverse probability of testing weights (sIPTW). Note: The motivation is to account for differences between tested and untested children when estimating RSV positivity.<ul><li>
<div class="bbcode_container">
	<div class="bbcode_description">Code:</div>
	<pre class="bbcode_code">logistic rsvtest $EN_vars_a1, vce(cluster child_ID)</pre>
</div></li>
<li>
<div class="bbcode_container">
	<div class="bbcode_description">Code:</div>
	<pre class="bbcode_code">predict rsvtest_a1</pre>
</div></li>
<li>
<div class="bbcode_container">
	<div class="bbcode_description">Code:</div>
	<pre class="bbcode_code">quietly sum rsvtest</pre>
</div></li>
<li>
<div class="bbcode_container">
	<div class="bbcode_description">Code:</div>
	<pre class="bbcode_code">local rsvtest_mean=r(mean)</pre>
</div></li>
<li>
<div class="bbcode_container">
	<div class="bbcode_description">Code:</div>
	<pre class="bbcode_code">replace rsvtest_sIPTW_a1=(1-`rsvtest_mean')/(1-rsvtest_p_a1) if rsvtest==0</pre>
</div></li>
<li>
<div class="bbcode_container">
	<div class="bbcode_description">Code:</div>
	<pre class="bbcode_code">replace rsvtest_sIPTW_a1=`rsvtest_mean'/rsvtest_p_a1 if rsvtest==1</pre>
</div></li>
</ul></li>
<li>Fit a weighted logistic regression model for RSV positivity among tested children.<ul><li>
<div class="bbcode_container">
	<div class="bbcode_description">Code:</div>
	<pre class="bbcode_code">logistic rsvpos $EN_vars_a1 if rsvtest==1 [pw=rsvtest_sIPTW_a1], vce(cluster b_ID)</pre>
</div></li>
<li>
<div class="bbcode_container">
	<div class="bbcode_description">Code:</div>
	<pre class="bbcode_code">predict rsvpos_p_a1</pre>
</div></li>
</ul></li>
<li>Evaluate model performance (e.g., sensitivity, specificity, Youden index) and choose a probability cutoff. For simplicity, I will use a cutoff of 0.5.<ul><li>
<div class="bbcode_container">
	<div class="bbcode_description">Code:</div>
	<pre class="bbcode_code">estat classification, cutoff(.5)</pre>
</div></li>
</ul></li>
<li>Apply the model to untested children and classify probable RSV-associated hospitalisations.<ul><li>
<div class="bbcode_container">
	<div class="bbcode_description">Code:</div>
	<pre class="bbcode_code">gen predictedRSV=0</pre>
</div></li>
<li>
<div class="bbcode_container">
	<div class="bbcode_description">Code:</div>
	<pre class="bbcode_code">replace predictedRSV=1 if rsvpos_p_a1&gt;=.5</pre>
</div></li>
</ul></li>
<li>Define the estimated &quot;true&quot; RSV burden as:<ul><li>
<div class="bbcode_container">
	<div class="bbcode_description">Code:</div>
	<pre class="bbcode_code">gen trueRSV=0</pre>
</div></li>
<li>
<div class="bbcode_container">
	<div class="bbcode_description">Code:</div>
	<pre class="bbcode_code">replace trueRSV=1 if rsvpos==1 | rsvtest==0 &amp; predictedRSV==1</pre>
</div></li>
</ul></li>
</ol>We also have plans to use the &quot;trueRSV&quot; variable for subsequent analyses.<br />
<br />
Questions:<ol class="decimal"><li>Does this seem like a reasonable framework for estimating the true burden of RSV in this setting?</li>
<li>Is the use of stabilised inverse probability of testing weights appropriate here, given that RSV status is only observed among tested children?</li>
<li>Are there alternative approaches (e.g., multiple imputation, two-stage selection models, semi-supervised learning, or other methods for outcome verification bias) that would be more appropriate?</li>
<li>Would it be preferable to estimate the RSV burden using the predicted probabilities directly (i.e., summing probabilities across untested children) rather than imposing a classification threshold such as 0.5?</li>
</ol>Kind regards,<br />
Damien]]></content:encoded>
			<category domain="https://www.statalist.org/forums/forum/general-stata-discussion/general">General</category>
			<dc:creator>Damien YP Foo</dc:creator>
			<guid isPermaLink="true">https://www.statalist.org/forums/forum/general-stata-discussion/general/1787446-prediction-models-and-stabilised-inverse-probability-of-treatment-weighting</guid>
		</item>
		<item>
			<title>2026 Stata Conference proceedings now available</title>
			<link>https://www.statalist.org/forums/forum/general-stata-discussion/general/1787438-2026-stata-conference-proceedings-now-available</link>
			<pubDate>Tue, 06 Oct 2026 13:10:35 GMT</pubDate>
			<description>Proceedings for presentations at the 2026 Stata Conference in Boston last week are now available on RePEc services:...</description>
			<content:encoded><![CDATA[Proceedings for presentations at the 2026 Stata Conference in Boston last week are now available on RePEc services: <a href="https://econpapers.repec.org/paper/bocusug26/" target="_blank">https://econpapers.repec.org/paper/bocusug26/</a>  and <a href="https://ideas.repec.org/s/boc/usug26.html" target="_blank">https://ideas.repec.org/s/boc/usug26.html</a><br />
 ]]></content:encoded>
			<category domain="https://www.statalist.org/forums/forum/general-stata-discussion/general">General</category>
			<dc:creator>KitBaum</dc:creator>
			<guid isPermaLink="true">https://www.statalist.org/forums/forum/general-stata-discussion/general/1787438-2026-stata-conference-proceedings-now-available</guid>
		</item>
		<item>
			<title>Plausible Exogeneity test on Control Function Instruments</title>
			<link>https://www.statalist.org/forums/forum/general-stata-discussion/general/1787436-plausible-exogeneity-test-on-control-function-instruments</link>
			<pubDate>Mon, 05 Oct 2026 16:37:32 GMT</pubDate>
			<description>I have used control function approach to test for endogeneity of my variables. Now I want to test the validity of my instrument using Plausible...</description>
			<content:encoded>I have used control function approach to test for endogeneity of my variables. Now I want to test the validity of my instrument using Plausible Exogeneity. Is there a way to do that in Stata?</content:encoded>
			<category domain="https://www.statalist.org/forums/forum/general-stata-discussion/general">General</category>
			<dc:creator>Daipayan Dhar</dc:creator>
			<guid isPermaLink="true">https://www.statalist.org/forums/forum/general-stata-discussion/general/1787436-plausible-exogeneity-test-on-control-function-instruments</guid>
		</item>
		<item>
			<title>SSC Archive access statistics</title>
			<link>https://www.statalist.org/forums/forum/general-stata-discussion/general/1787432-ssc-archive-access-statistics</link>
			<pubDate>Sun, 04 Oct 2026 23:58:17 GMT</pubDate>
			<description>Statistics for September 2026 access to the SSC Archive are now available at  http://repec.org/docs/sscAuthors.html  shows the oervall ranking of...</description>
			<content:encoded><![CDATA[Statistics for September 2026 access to the SSC Archive are now available at  <a href="http://repec.org/docs/sscAuthors.html" target="_blank">http://repec.org/docs/sscAuthors.html</a>  shows the oervall ranking of almost 1600 authors of community-contributed software for the last three months. The sec hot command gives rankings for the latest month, and can provide rankings for more packages or specific by author.  There are presently 4020 Stata packages on the SSC Archive, including six that have been added this month.]]></content:encoded>
			<category domain="https://www.statalist.org/forums/forum/general-stata-discussion/general">General</category>
			<dc:creator>KitBaum</dc:creator>
			<guid isPermaLink="true">https://www.statalist.org/forums/forum/general-stata-discussion/general/1787432-ssc-archive-access-statistics</guid>
		</item>
		<item>
			<title>wrong SE in dsregress with cluster() option</title>
			<link>https://www.statalist.org/forums/forum/general-stata-discussion/general/1787426-wrong-se-in-dsregress-with-cluster-option</link>
			<pubDate>Sat, 03 Oct 2026 15:54:00 GMT</pubDate>
			<description>Hi statalisters,  
 
it looks like when you run the following model: 
 
 
dsregress Y T, controls(X1 X2...) cluster (clustervar)  
 
 
dsregress does...</description>
			<content:encoded><![CDATA[Hi statalisters, <br />
<br />
it looks like when you run the following model:<br />

<div class="bbcode_container">
	<div class="bbcode_description">Code:</div>
	<pre class="bbcode_code">dsregress Y T, controls(X1 X2...) cluster (clustervar)</pre>
</div>dsregress does not generate an error and simply provide the double lasso estimator without correcting for clustering. <br />
<br />
The correct code should be (as you probably know):  <br />

<div class="bbcode_container">
	<div class="bbcode_description">Code:</div>
	<pre class="bbcode_code">dsregress Y T, controls(<b>X1 X2...) vce(cluster clustervar) </b></pre>
</div>Yet, why does stata authorize the first coding? In a classical regression reg Y T, cluster(clustervar) and  regression reg Y T, vce(cluster clustervar) provide the same SE. <br />
<br />
I think this should not be allowed since it is a source of error. <br />
<br />
What do you think? <br />
 ]]></content:encoded>
			<category domain="https://www.statalist.org/forums/forum/general-stata-discussion/general">General</category>
			<dc:creator>Adrien Bouguen</dc:creator>
			<guid isPermaLink="true">https://www.statalist.org/forums/forum/general-stata-discussion/general/1787426-wrong-se-in-dsregress-with-cluster-option</guid>
		</item>
		<item>
			<title>Help with using Monte Carlo Simulation to do sensitivity power analysis on a mixed multilevel model</title>
			<link>https://www.statalist.org/forums/forum/general-stata-discussion/general/1787421-help-with-using-monte-carlo-simulation-to-do-sensitivity-power-analysis-on-a-mixed-multilevel-model</link>
			<pubDate>Fri, 02 Oct 2026 02:42:07 GMT</pubDate>
			<description>Hi,  
I am a psychology honours student who is using a mixed multilevel model for one of my research questions. I am considering using Monte Carlo...</description>
			<content:encoded><![CDATA[Hi, <br />
I am a psychology honours student who is using a mixed multilevel model for one of my research questions. I am considering using Monte Carlo simulation to do a sensitivity power analysis. However, we have not learn this at uni, and I am not sure if I am doing it correctly. Would someone be able to check if my coding is correct? It seems to run fine on my version of Stata 17, but as I am unfamiliar with it I just wanted to check if I have done the coding correctly. My variables are<br />
pid: participant ID (level 2)<br />
agegrp: age group, 1 = older, 2 = younger (older is reference level)<br />
cval: cue valence, 1 = negative, 2 = neutral, 3 = positive<br />
avg_rating_cue:  affective appraisal rating (DV), averaged within each cue valence condition<br />
prop_spec_val_c: proportion of specific memories in each valence condition (predictor, grand centred)<br />
<br />
<b>Step 1: Estimate inputs from the real data</b><br />
* SD components of the predictor (valence main effect removed)<br />
mixed prop_spec_val_c i.cval || pid:, reml<br />
scalar sd_spec_b   = exp(_b[lns1_1_1:_cons])     // between-person SD<br />
scalar sd_spec_w   = exp(_b[lnsig_e:_cons])      // within-person SD<br />
scalar sd_spec_tot = sqrt(sd_spec_b^2 + sd_spec_w^2)<br />
<br />
* Outcome model (observed)<br />
mixed avg_rating_cue i.agegrp##c.prop_spec_val_c i.cval || pid:, reml<br />
scalar sd_intercept = exp(_b[lns1_1_1:_cons])    // random-intercept SD<br />
scalar sd_resid     = exp(_b[lnsig_e:_cons])     // residual SD<br />
scalar b_cons       = _b[_cons]<br />
scalar b_agegrp     = _b[2.agegrp]<br />
scalar b_spec       = _b[c.prop_spec_val_c]      // slope in the reference group (older)<br />
<br />
<b>Step 2: simulate one dataset and test the interaction</b><br />
capture program drop sim_rq3<br />
program sim_rq3, rclass<br />
    version 16<br />
    syntax, delta(real) n1(integer) n2(integer) nobs(integer) [alpha(real 0.05)]<br />
    local ntot = `n1' + `n2'<br />
    local reject = .<br />
    local conv   = 0<br />
    quietly {<br />
        drop _all<br />
        set obs `ntot'<br />
        generate long pid = _n<br />
        generate byte agegrp = cond(_n &lt;= `n1', 1, 2)   // 1 = Older, 2 = Younger<br />
<br />
        generate double u_i = rnormal(0, sd_intercept)  // random intercept<br />
        generate double s_i = rnormal(0, sd_spec_b)     // person's mean specificity<br />
        expand `nobs'                                   // 3 observations per person<br />
        generate double spec = s_i + rnormal(0, sd_spec_w)<br />
<br />
        * delta = true slope difference (added to group 1)<br />
        generate double slope = cond(agegrp==1, b_spec + `delta', b_spec)<br />
        generate double rating = b_cons + b_agegrp*(agegrp==2) ///<br />
            + slope*spec + u_i + rnormal(0, sd_resid)<br />
<br />
        capture mixed rating i.agegrp##c.spec || pid:, reml iter(200)<br />
        if _rc == 0 {<br />
            if e(converged) == 1 {<br />
                local conv = 1<br />
                testparm i.agegrp#c.spec<br />
                local p = chi2tail(r(df), r(chi2))      // Wald test<br />
                local reject = (`p' &lt; `alpha')<br />
            }<br />
        }<br />
    }<br />
    return scalar reject = `reject'<br />
    return scalar conv   = `conv'<br />
end<br />
<br />
<b>Step 3: power curve</b><br />
local n_older = 34<br />
local n_younger = 36<br />
local nobs = 3<br />
local reps = 1000<br />
local seed = 12345<br />
<br />
tempname results<br />
postfile `results' double delta double delta_per_sd double power double conv_rate ///<br />
    using rq3_power_curve.dta, replace<br />
<br />
foreach d of numlist 0 1.0(0.1)1.4 {<br />
    quietly simulate reject=r(reject) conv=r(conv), reps(`reps') seed(`seed') nodots: ///<br />
        sim_rq3, delta(`d') n1(`n_older') n2(`n_younger') nobs(`nobs')<br />
    quietly summarize reject<br />
    local pwr = r(mean)<br />
    quietly summarize conv<br />
    local cr = r(mean)<br />
    post `results' (`d') (`d'*sd_spec_tot) (`pwr') (`cr')<br />
}<br />
postclose `results'<br />
<br />
Thank you for your help.]]></content:encoded>
			<category domain="https://www.statalist.org/forums/forum/general-stata-discussion/general">General</category>
			<dc:creator>Apostolos Ng</dc:creator>
			<guid isPermaLink="true">https://www.statalist.org/forums/forum/general-stata-discussion/general/1787421-help-with-using-monte-carlo-simulation-to-do-sensitivity-power-analysis-on-a-mixed-multilevel-model</guid>
		</item>
		<item>
			<title>get AMEs by combining -margins- with some sort of parallel process?</title>
			<link>https://www.statalist.org/forums/forum/general-stata-discussion/general/1787418-get-ames-by-combining-margins-with-some-sort-of-parallel-process</link>
			<pubDate>Thu, 01 Oct 2026 16:55:51 GMT</pubDate>
			<description><![CDATA[Dear Statalist: 
 
I am running a massive multinomial logistic regression using restricted data on an organization's servers. I'm trying to get...]]></description>
			<content:encoded><![CDATA[Dear Statalist:<br />
<br />
I am running a massive multinomial logistic regression using restricted data on an organization's servers. I'm trying to get average marginal effects.<br />
<br />
This is not my machine, I have to connect to it remotely and use their virtual machine. This means that upgrading to a Stata / MP that can use 16 (or 32, or 64) cores is not an option (unfortunately).<br />
<br />
I had an idea to use the user-written package -parallel- by Yon and Quistorff to try and speed it up. I do not know if this is possible, but if it is theoretically possible I would like to try.<br />
<br />
Some context:<br />
This multinomial logistic regression has 60 outcomes, and the sample I'm running it on has over six million respondents. The multinomial logistic regression usually finishes in about 12-15 hours, and if I get predicted values using -margins- with no standard errors at all (-nose- option) it usually takes about 45 hours. I am currently trying to get predicted values using margins with standard errors and it is definitely using every last drop of RAM to do it--currently about about 110 gigabytes of RAM. I suppose it might even take more if I had the RAM to spare. I'm also not sure when it will finish. Maybe tomorrow, maybe Monday, maybe never.<br />
<br />
However, something very different happened when I tried to get average marginal effects (AMEs) using -margins- and getting standard errors. The process started off by using up huge chunks of processing power (between 15% and 50% of the available CPU) and about 3 gigabytes of RAM for about a day. It then dropped down to about 5% to 20% of CPU, but increased to using about 6 gigabytes of RAM. And it kept doing this for about a week before I finally killed the process.<br />
<br />
I suppose I can just set it up to run and come back and check it every week until it's done, but I'm not sure it will ever finish that way. I'm also not sure why it's not using more available RAM. I'm hoping that maybe there's a way to parallelize the AMEs further, chopping up my RAM into 6 gigabyte chunks across multiple instances. But that depends very much on how -margins- calculates the AME, and it is possible I am misunderstanding that part.<br />
<br />
Has anyone tried doing something like this with -parallel-? Or anything else? If I'm patient enough eventually I can get predicted values without the SEs, but getting SEs for the AMEs seems to require something stronger.<br />
<br />
Thanks for your time,<br />
Jonathan]]></content:encoded>
			<category domain="https://www.statalist.org/forums/forum/general-stata-discussion/general">General</category>
			<dc:creator>Jonathan Horowitz</dc:creator>
			<guid isPermaLink="true">https://www.statalist.org/forums/forum/general-stata-discussion/general/1787418-get-ames-by-combining-margins-with-some-sort-of-parallel-process</guid>
		</item>
		<item>
			<title>SSC activity, September 2026</title>
			<link>https://www.statalist.org/forums/forum/general-stata-discussion/general/1787411-ssc-activity-september-2026</link>
			<pubDate>Thu, 01 Oct 2026 02:23:19 GMT</pubDate>
			<description>During the last month, 28 packages have been added to the SSC Archive, and 42 packages have been updated, for a total of 4,014 packages. 
 
At this...</description>
			<content:encoded><![CDATA[During the last month, 28 packages have been added to the SSC Archive, and 42 packages have been updated, for a total of 4,014 packages.<br />
<br />
At this week's Stata Conference in Boston, I will present details of the SSC-Next Generation project, which will transform the archive. One aspect of that project is already implemented: the ssc2 package (sec describe ssc2) provides access to earlier versions of community-contributed packages to enhance reproducibility. The Stata Conference RePEc series will share aspects of the SSC-NG project in the next few days.]]></content:encoded>
			<category domain="https://www.statalist.org/forums/forum/general-stata-discussion/general">General</category>
			<dc:creator>KitBaum</dc:creator>
			<guid isPermaLink="true">https://www.statalist.org/forums/forum/general-stata-discussion/general/1787411-ssc-activity-september-2026</guid>
		</item>
		<item>
			<title>Stata Sociology Virtual Symposium 2026</title>
			<link>https://www.statalist.org/forums/forum/general-stata-discussion/general/1787407-stata-sociology-virtual-symposium-2026</link>
			<pubDate>Wed, 30 Sep 2026 08:21:37 GMT</pubDate>
			<description>The 2026 Stata Sociology Virtual Symposium is a meeting of researchers in sociology from around the world discussing current theory and applied...</description>
			<content:encoded><![CDATA[The 2026 Stata Sociology Virtual Symposium is a meeting of researchers in sociology from around the world discussing current theory and applied methods using Stata. The proceedings consist of invited talks by top Stata users in a virtual platform that allows you to experience this one-day event from wherever you are. The symposium is organized by StataNordic and Metrika Consulting.<br />
<br />
<br />
<a href="filedata/fetch?filedataid=1787408">Array </a><br />
<br />
<br />
Program and Registration available at: <a href="https://www.statanordic.com/stata-sociology-virtual-symposium26.html" target="_blank">https://www.statanordic.com/stata-so...mposium26.html</a>]]></content:encoded>
			<category domain="https://www.statalist.org/forums/forum/general-stata-discussion/general">General</category>
			<dc:creator>Samuel Mossberg</dc:creator>
			<guid isPermaLink="true">https://www.statalist.org/forums/forum/general-stata-discussion/general/1787407-stata-sociology-virtual-symposium-2026</guid>
		</item>
		<item>
			<title>Winsorization of variables using winsor2</title>
			<link>https://www.statalist.org/forums/forum/general-stata-discussion/general/1787403-winsorization-of-variables-using-winsor2</link>
			<pubDate>Tue, 29 Sep 2026 12:04:57 GMT</pubDate>
			<description>I winsorized 25 variables simultaneously using the winsor2 command in Stata. However, I noticed that the first variable in the variable list does not...</description>
			<content:encoded><![CDATA[I winsorized 25 variables simultaneously using the winsor2 command in Stata. However, I noticed that the first variable in the variable list does not appear to have been winsorized correctly and contains some unusual values after running the command. In contrast, the remaining variables appear to have been winsorized as expected.<br />
<br />
Is this a known issue with winsor2, or could it be related to the way the command is specified? Has anyone encountered a situation where the first variable in a winsor2 command is treated differently from the others?<br />
<br />
Any guidance on how to diagnose or resolve this issue would be greatly appreciated.<br />
 ]]></content:encoded>
			<category domain="https://www.statalist.org/forums/forum/general-stata-discussion/general">General</category>
			<dc:creator>Bose Sudipta</dc:creator>
			<guid isPermaLink="true">https://www.statalist.org/forums/forum/general-stata-discussion/general/1787403-winsorization-of-variables-using-winsor2</guid>
		</item>
		<item>
			<title>hmerge: a prototype hash-based -merge- (m:1, 1:1) that skips the master sort, benchmarks and request for testers</title>
			<link>https://www.statalist.org/forums/forum/general-stata-discussion/general/1787392-hmerge-a-prototype-hash-based-merge-m-1-1-1-that-skips-the-master-sort-benchmarks-and-request-for-testers</link>
			<pubDate>Fri, 25 Sep 2026 03:00:35 GMT</pubDate>
			<description>It occurred to me today that once upon a time I remember gtools having a long to-do list of speedups the author wanted to implement but never did,...</description>
			<content:encoded><![CDATA[It occurred to me today that once upon a time I remember gtools having a long to-do list of speedups the author wanted to implement but never did, and that there were probably plenty of speedups we might get if we turned agents onto the problem of extending the work of gtools or finding other speedups. When I asked it to take a look at everything, it came up with a speedup to merge. What follows below is its writeup of those results, and I have confirmed the code runs on my own computer.<br />
<br />
I have a prototype command, hmerge, that does merge m:1 and merge 1:1 with a hash table instead of sorting the master data. With the same syntax as merge and identical results, it runs about 8 times faster than merge on 10 million master observations with one integer key and up to 100,000 using observations (5.4 to 13.6 times across the key types I tried), when you don't need the result re-sorted by the merge key. The gain is smaller with very large using files. I'd love feedback, and especially people willing to try to break it.<br />
<br />
Setup for everything below: Stata/MP 17.0 (2-core licence) on a MacBook Pro (Apple M5 Pro, 24 GB, macOS 26.5). hmerge is a C plugin plus an ado wrapper, so for now it only exists as a macOS arm64 build. Times are seconds, median of 3 runs with a warm file cache, end to end (master in memory, using file on disk, merged data in memory).<br />
<br />
<b>Why it's faster</b><br />
<br />
merge sorts the whole master dataset by the key, moving every variable, before a linear sort-merge join. That sort is nearly all of the cost. With a master already sorted by the key, merge m:1 on 10 million observations takes 0.283s instead of 2.015s. hmerge instead reads the using data into a scratch frame, builds an index on its keys (a direct lookup table for one integer key over a compact range, otherwise a hash table), then looks up each master row where it already sits. The attached figure (hmerge_sortmerge_vs_hashjoin.png) shows the two approaches side by side. Hash collisions can't produce wrong matches, because a match is always confirmed by comparing the full key values.<br />
<br />
Headline numbers<br />
<br />
merge m:1 on one integer key, 10 million master observations in random order, 3 using variables:<br />
<br />

<div class="bbcode_container">
	<div class="bbcode_description">Code:</div>
	<pre class="bbcode_code">using obs (J)    merge   hmerge   hmerge,sort   fmerge   frames   merge/hmerge
          100    1.924    0.231         1.929    1.018    7.070          8.3x
      100,000    2.015    0.259         2.028    1.002    7.447          7.8x
    5,000,000    2.299    1.204         3.370    2.910   11.449          1.9x</pre>
</div>fmerge is from ftools (SSC), and frames means frame create + use + frlink + frget. At 100,000 using observations the speedup is 5.4x to 13.6x across key types (13.6x for a str12 key, 9.8x for a mixed string and numeric key, which fmerge doesn't accept).<br />
<br />
<b>Where it doesn't help</b><br />
<br />

<div class="bbcode_container">
	<div class="bbcode_description">Code:</div>
	<pre class="bbcode_code">case (10 million obs, J = 100,000)                          merge   hmerge   ratio
master already sorted by the key                            0.283    0.239    1.2x
merge, then bysort on the key right after                   2.138    2.270    0.94x
panel sorted by person-year that has to keep that order     4.373    0.272    16.1x
1:1 on a unique integer key (J = 10 million)                2.377    1.602    1.5x</pre>
</div>So the gain is skipping a sort you don't need. If you want the result sorted by the merge key anyway, merge is as fast or slightly faster. One thing about order: after merge m:1 the master observations are in key order, but Stata no longer treats the data as sorted, so a -by key:- afterwards needs a sort with either command. hmerge leaves the master observations in their original order. Both commands add using-only observations at the end, in key order.<br />
<br />
<b>A public benchmark</b><br />
<br />
To check that this isn't an artifact of my own simulated data, I ran the join task from the H2O.ai/DuckDB Labs &quot;database-like ops&quot; benchmark (<a href="https://github.com/duckdblabs/db-benchmark" target="_blank">https://github.com/duckdblabs/db-benchmark</a>, commit 3b074bc, with results for other software at <a href="https://duckdblabs.github.io/db-benchmark/" target="_blank">https://duckdblabs.github.io/db-benchmark/</a>). Its generator makes a 10 million observation table and three using tables. I mapped its five questions to merge with keepusing(v2) and added R's data.table on the same machine, limited to 2 threads to match Stata/MP:<br />
<br />

<div class="bbcode_container">
	<div class="bbcode_description">Code:</div>
	<pre class="bbcode_code">question                         merge   hmerge   fmerge   data.table(2 thr)
q1 small inner (J=10)            1.681    0.323    0.878               0.104
q2 medium inner (J=10,000)       1.959    0.354    0.919               0.178
q3 medium outer (keep master)    1.861    0.163    0.719               0.222
q4 medium inner, string key      3.282    0.485    3.495               0.151
q5 big inner, 1:1 (J=10M)        4.182    1.621    6.635               1.103</pre>
</div>All implementations return the same number of observations, and hmerge's result is identical to merge's on all five. The Stata times include reading the using .dta from disk, while data.table already has its tables in memory.<br />
<br />
<b>Correctness</b><br />
<br />
- 130 differential test cases compare hmerge with merge on identical inputs: values, _merge, variable order, storage types (including merge's type promotion), formats, variable and value labels, error codes, the data left in memory after an error, and whether any sort flag left on the data is true.<br />
- The tests also pass with the plugin compiled in a checking mode (UndefinedBehaviorSanitizer) that halts on any undefined behavior in the C code.<br />
<br />
Other differences from merge: variable notes aren't copied, and r() holds r(path), which records how hmerge did the join, instead of what merge leaves there. For 1:m and m:m merges, the update, replace and force options, and strL variables, hmerge hands the job to merge and says so.<br />
<br />
<b>Limitations</b><br />
<br />
It's only been run on Stata 17 with 2 cores on a Mac. Stata 18/19, Windows, Linux and more cores are all untested. To me the shakiest part is that the design relies on plugin memory persisting between two plugin calls, which works but isn't documented in the plugin interface, so a future Stata could break it (it is guarded so a mismatch errors out rather than using stale data). At a very large using file hmerge uses somewhat more memory than merge (594 vs 498 MB above the data itself with 5 million using observations).<br />
<br />
<b>What would help most</b><br />
<br />
1. If you can build and run it on Windows or Linux, or on Stata 18/19 or more cores, and report timings.<br />
2. Any case where hmerge and merge give different results.<br />
<br />
Code, tests, benchmark do-files and a longer write-up: <a href="https://github.com/clibassi/hmerge" target="_blank">https://github.com/clibassi/hmerge</a> (install: net install hmerge, from(&quot;https://raw.githubusercontent.com/clibassi/hmerge/main/&quot;)). A short demo do-file that builds a 10 million observation person-year panel and times both commands is in demo/hmerge_speedup_demo.do.<br />
<br />
The index idea for integer keys follows gtools by Mauricio Cáceres Bravo (SSC, <a href="https://github.com/mcaceresb/stata-gtools" target="_blank">https://github.com/mcaceresb/stata-gtools</a>), and ftools by Sergio Correia (SSC, <a href="https://github.com/sergiocorreia/ftools" target="_blank">https://github.com/sergiocorreia/ftools</a>) was the obvious comparison. Claude built this with a little nudging from me.<br />
<br />
If you see mistakes or notice any crazy choices, that would be super helpful. Thanks!<br />
<br />
CJ Libassi<br />
<br />
     <a href="filedata/fetch?filedataid=1787393">Array </a>]]></content:encoded>
			<category domain="https://www.statalist.org/forums/forum/general-stata-discussion/general">General</category>
			<dc:creator>CJ Libassi</dc:creator>
			<guid isPermaLink="true">https://www.statalist.org/forums/forum/general-stata-discussion/general/1787392-hmerge-a-prototype-hash-based-merge-m-1-1-1-that-skips-the-master-sort-benchmarks-and-request-for-testers</guid>
		</item>
		<item>
			<title>Reducing data set</title>
			<link>https://www.statalist.org/forums/forum/general-stata-discussion/general/1787390-reducing-data-set</link>
			<pubDate>Thu, 24 Sep 2026 20:22:51 GMT</pubDate>
			<description><![CDATA[Is there a way to exclude cases from a dataset if they do not have sufficient observations? 
 
I wrote: 
 
keep house if _n(VCF0900c &gt;= 45) 
 
but it...]]></description>
			<content:encoded><![CDATA[Is there a way to exclude cases from a dataset if they do not have sufficient observations?<br />
<br />
I wrote:<br />
<br />
keep house if _n(VCF0900c &gt;= 45)<br />
<br />
but it didn't work  Is there a solution to this problem?  Thanks for any help.<br />
 ]]></content:encoded>
			<category domain="https://www.statalist.org/forums/forum/general-stata-discussion/general">General</category>
			<dc:creator>Euslaner</dc:creator>
			<guid isPermaLink="true">https://www.statalist.org/forums/forum/general-stata-discussion/general/1787390-reducing-data-set</guid>
		</item>
		<item>
			<title>New Stata command: xtie2 – Interactive panel data estimation with automatic pairwise interactions</title>
			<link>https://www.statalist.org/forums/forum/general-stata-discussion/general/1787386-new-stata-command-xtie2-–-interactive-panel-data-estimation-with-automatic-pairwise-interactions</link>
			<pubDate>Thu, 24 Sep 2026 05:44:01 GMT</pubDate>
			<description>Dear Statalist members, 
 
I would like to introduce xtie2, a new user-written Stata command for fixed-effects panel estimation with automatic...</description>
			<content:encoded><![CDATA[Dear Statalist members,<br />
<br />
I would like to introduce <b>xtie2</b>, a new user-written Stata command for fixed-effects panel estimation with automatic pairwise interactions.<br />
<br />
The command automatically generates all pairwise interactions among the explanatory variables and performs post-estimation analysis for the primary interaction defined by the first two explanatory variables. It reports the interaction coefficient and significance test, calculates marginal effects, evaluates them at selected moderator values, and computes marginal-effect zero-crossing critical values. It also checks whether these critical values fall within the observed sample range.<br />
<br />
The reported critical values should be interpreted as <b>marginal-effect zero-crossing values</b>, not as Hansen-type panel threshold estimates.<br />
<br />
The package is currently available from GitHub and can be installed in Stata using:<br />
<br />
net install xtie2, from(&quot;https://raw.githubusercontent.com/zehrayalnizz41-cmd/xtie2/main/&quot;)<br />
<br />
After installation:<br />
<br />
help xtie2<br />
<br />
Example:<br />
<br />
xtie2 unemployment energy_dependency energy_price_index inflation gdp_growth, robust<br />
<br />
GitHub repository:<br />
<a href="https://github.com/zehrayalnizz41-cmd/xtie2" target="_blank">https://github.com/zehrayalnizz41-cmd/xtie2</a><br />
<br />
I would be grateful for comments, suggestions, and bug reports from Stata users.<br />
<br />
Best regards,<br />
<b>Dr. Zehra Yalnız</b><br />
Kocaeli University, Türkiye<br />
ORCID: 0000-0003-2633-2022<br />
 ]]></content:encoded>
			<category domain="https://www.statalist.org/forums/forum/general-stata-discussion/general">General</category>
			<dc:creator>Zehra Yalniz</dc:creator>
			<guid isPermaLink="true">https://www.statalist.org/forums/forum/general-stata-discussion/general/1787386-new-stata-command-xtie2-–-interactive-panel-data-estimation-with-automatic-pairwise-interactions</guid>
		</item>
		<item>
			<title>relative survival poisson model</title>
			<link>https://www.statalist.org/forums/forum/general-stata-discussion/general/1787384-relative-survival-poisson-model</link>
			<pubDate>Wed, 23 Sep 2026 23:07:18 GMT</pubDate>
			<description><![CDATA[HI Statalist, 
I am trying to model rs using a poisson model as per Royston and Lambert's book Flexible Paremetric Survival Analysis using Stata...]]></description>
			<content:encoded><![CDATA[HI Statalist,<br />
I am trying to model rs using a poisson model as per Royston and Lambert's book <i>Flexible Paremetric Survival Analysis using Stata (2011)</i>. I have found the mean/variance is not equal. Does this matter? Is there a way of correcting for this in rs models?<br />
Thanks,<br />
Matt]]></content:encoded>
			<category domain="https://www.statalist.org/forums/forum/general-stata-discussion/general">General</category>
			<dc:creator>Matt Piercy</dc:creator>
			<guid isPermaLink="true">https://www.statalist.org/forums/forum/general-stata-discussion/general/1787384-relative-survival-poisson-model</guid>
		</item>
	</channel>
</rss>
