Hi all,
I’d appreciate some thoughts on the following methodological question:
I’m working on a proposal to evaluate a physical activity intervention in primary schools, using a cluster non-randomised design (schools as clusters). We will have baseline and follow-up data for treatment and control schools (12 per arm), so this is a standard two-group, two-time-point setting.
The intervention is co-designed and implemented with schools and local communities, so assignment to treatment is likely driven by both observed and unobserved school characteristics (e.g. leadership, motivation, capacity, local need). These characteristics may also directly affect outcomes, making selection bias a key concern.
Within the team, one option proposed by the statistician is a linear mixed model (LMM) to account for clustering and baseline differences. An alternative proposed by me is a difference-in-differences approach. While both approaches may yield similar point estimates here, my main concern is the identification assumptions rather than estimation per se. DiD relies on a parallel trends assumption and allows time-invariant unobserved factors to be correlated with treatment assignment. I'm not an expert in statistics, but my understanding is that causal interpretation of LMMs in non-randomised settings may implicitly rely on stronger conditional independence/exchangeability assumptions, i.e. that school characteristics are not related to treatment assignment once modelled, which seem harder to justify in this context.
I’d be very grateful for any thoughts on this, or references discussing DiD versus mixed models from a causal inference perspective in non-randomised cluster settings.
Many thanks in advance!
I’d appreciate some thoughts on the following methodological question:
I’m working on a proposal to evaluate a physical activity intervention in primary schools, using a cluster non-randomised design (schools as clusters). We will have baseline and follow-up data for treatment and control schools (12 per arm), so this is a standard two-group, two-time-point setting.
The intervention is co-designed and implemented with schools and local communities, so assignment to treatment is likely driven by both observed and unobserved school characteristics (e.g. leadership, motivation, capacity, local need). These characteristics may also directly affect outcomes, making selection bias a key concern.
Within the team, one option proposed by the statistician is a linear mixed model (LMM) to account for clustering and baseline differences. An alternative proposed by me is a difference-in-differences approach. While both approaches may yield similar point estimates here, my main concern is the identification assumptions rather than estimation per se. DiD relies on a parallel trends assumption and allows time-invariant unobserved factors to be correlated with treatment assignment. I'm not an expert in statistics, but my understanding is that causal interpretation of LMMs in non-randomised settings may implicitly rely on stronger conditional independence/exchangeability assumptions, i.e. that school characteristics are not related to treatment assignment once modelled, which seem harder to justify in this context.
I’d be very grateful for any thoughts on this, or references discussing DiD versus mixed models from a causal inference perspective in non-randomised cluster settings.
Many thanks in advance!

Comment