Announcement

Collapse
No announcement yet.
X
  • Filter
  • Time
  • Show
Clear All
new posts

  • DiD vs mixed models in non-randomised studies

    Hi all,

    I’d appreciate some thoughts on the following methodological question:

    I’m working on a proposal to evaluate a physical activity intervention in primary schools, using a cluster non-randomised design (schools as clusters). We will have baseline and follow-up data for treatment and control schools (12 per arm), so this is a standard two-group, two-time-point setting.

    The intervention is co-designed and implemented with schools and local communities, so assignment to treatment is likely driven by both observed and unobserved school characteristics (e.g. leadership, motivation, capacity, local need). These characteristics may also directly affect outcomes, making selection bias a key concern.

    Within the team, one option proposed by the statistician is a linear mixed model (LMM) to account for clustering and baseline differences. An alternative proposed by me is a difference-in-differences approach. While both approaches may yield similar point estimates here, my main concern is the identification assumptions rather than estimation per se. DiD relies on a parallel trends assumption and allows time-invariant unobserved factors to be correlated with treatment assignment. I'm not an expert in statistics, but my understanding is that causal interpretation of LMMs in non-randomised settings may implicitly rely on stronger conditional independence/exchangeability assumptions, i.e. that school characteristics are not related to treatment assignment once modelled, which seem harder to justify in this context.

    I’d be very grateful for any thoughts on this, or references discussing DiD versus mixed models from a causal inference perspective in non-randomised cluster settings.

    Many thanks in advance!

  • #2
    Hello, Alfredo.

    It seems like your case might be similar to non-cluster, non-randomized designs, where ANCOVA models aren’t recommended. In other words, ANCOVA models (which adjust for baseline values) are generally not used in observational studies (see links below). Since your project doesn’t involve randomization, as you mentioned, this situation is similar to the advice against using the ANCOVA model (but in your case the adjustment would be via the mixed-effects model).

    Links:

    https://www.sciencedirect.com/scienc...95435606000813
    https://tidsskriftet.no/en/2021/11/m...tional-studies

    Comment


    • #3
      Hi Alfedo,

      You are right that the school characteristics that lead a school to opt into your intervention might confound the results in the mixed effects model (and could potentially confound the DiD if certain assumptions are not met). Both approaches have their advantages. The mixed effects model gives you a powerful system to model the relationship between first and second order units from theory, whereas the DiD model is good for controlling for unobserved confounders. Ideally I would want to take a random sample of schools from the population of interest, randomly assign those schools to the intervention and control group, then use a mixed effects model for the best of both worlds, but it doesn't sound like that's practical here.

      I'd say selection of the model depends on what you want to do. If you want to rigorously decide whether your intervention works, then I'd go with the DiD. Just keep in mind that your audience in the education lit might be more used to seeing a mixed effects model (depending on the field), so a reviewer might ask questions about within-cluster effects that the DiD is not well suited to answer. In the DiD, would you look for student-level differences and cluster standard errors by school, even though the intervention is at the school level? Or would you aggregate student-level data to the school? Both are imperfect models of the clustered data. If you want to tell a compelling theory driven story about the way schools and students interact, you probably want the mixed model. If you want to do rigorous causal inference with mixed methods, you need random assignment to conditions. So there are tradeoffs.

      If I were you and your team, I would just do both and write two papers. DiD for the causal inference "does my intervention work" question, and a rich and theoretically driven mixed effects paper that gets into the details about how schools influence students both with and without intervention, how students vary within schools, and potentially the way students influence schools. The latter paper will require you to anticipate and measure important school level confounders, but it sounds like a more interesting read to me. The former is a more rigorous way to control for unobserved effects in the absence of random assignment to conditions, and might make for a better case if you want to say "my intervention works". Two great papers here.

      I take it that Tiago mentioned the ANCOVA model because that's what you see in Psychology and experimental design broadly, but I agree that it isn't appropriate here. Should give you results similar/equivalent to an OLS regression with an interaction between schools and students, I believe.
      Last edited by Daniel Schaefer; 31 Dec 2025, 10:50.

      Comment


      • #4
        I think the introduction of mixed models here is a bit of a distraction. To me, the choices are, (1) doing a 2 x 2 diff-in-diffs — which is the same as using the difference in Y across time in D, X, and D*(X - X1bar), where X1bar is the average of the covariates over treated units. The second choice is adding the first period outcome, also interactied with treatment D after centering. Under the DiD assumptions, adding lagged Y is generally over controlling. But one could just assume unconfoundedness conditional on X and Y_1. The latter is more efficient under an RCT in general. Sometimes people do both and argue they bound the average treatment effect on the treated.

        Comment


        • #5
          Originally posted by Jeff Wooldridge View Post
          I think the introduction of mixed models here is a bit of a distraction. To me, the choices are, (1) doing a 2 x 2 diff-in-diffs — which is the same as using the difference in Y across time in D, X, and D*(X - X1bar), where X1bar is the average of the covariates over treated units. The second choice is adding the first period outcome, also interactied with treatment D after centering. Under the DiD assumptions, adding lagged Y is generally over controlling. But one could just assume unconfoundedness conditional on X and Y_1. The latter is more efficient under an RCT in general. Sometimes people do both and argue they bound the average treatment effect on the treated.
          The discussion regarding the mixed-effects model is highly useful. It presents a basic valid argument, suggesting that mixed-effects modeling may not be appropriate for the design. Therefore, by simple process of elimination, DiD emerges as a strong choice.
          Last edited by Tiago Pereira; 01 Jan 2026, 14:35.

          Comment


          • #6
            But whether you use a DiD approach or include baseline outcomes in the equation, you have to account for the clustered assignment at the school level. (I assume the outcome is measured at the student level.) So setting it up as DiD versus mixed conflates two very different considerations. The simplest thing is to cluster standard errors at the school level after OLS estimation. Or, one could apple a mixed model in either case.

            Comment


            • #7
              I agree with Jeff. I should have written "suggesting that mixed-effects modeling with adjustment for baseline differences may not be appropriate for the study design".

              Comment


              • #8
                Many thanks to everyone for the very helpful discussion. I really appreciate the time and clarity of the feedback.

                Just to clarify, my objective is to estimate the causal effect of the programme. The intervention is implemented at the school level, but outcomes are measured at the student level, so I fully agree that clustering at the school level must be accounted for (and I was indeed considering clustered standard errors).

                My initial thought was to propose a standard 2×2 DiD, but I also appreciate the suggestion of considering a specification that includes the first-period outcome. Thanks, Jeff!

                The key issue motivating my question was not estimation per se, but the identification assumptions required to interpret estimates causally. In particular, my query was whether a linear mixed model in a non-randomised cluster setting carries implicit identification assumptions when adjusting (or not) for baseline values (thanks for the references, Tiago!). I am aware that LMMs are a natural choice in randomised cluster trials, but I was less clear on how to think about their identifying assumptions in non-randomised settings, where selection into treatment is likely, when adjusting (or not) for baseline values.

                Thanks again for all the insights and suggestions. This has been very helpful!

                Comment

                Working...
                X