Dear all,
I am having issues with a sample size calculation and would appreciate your advice.
Here is the scenario: I have mothers who are taking medications for an encephalopathy and who are also breastfeeding.
These drugs can be transferred into breast milk and then into the infants’ bloodstream.
I need to quantify the correlation between the amount of drug in breast milk and the amount of drug in the infants’ blood.
If we assumed that each mother was taking only one drug, and that only one drug was being studied, the problem would be straightforward (e.g., assuming a a priori Pearson correlation coefficient of 0.4, we would need roughly 40–50 mother–infant pairs).
However, in this study each mother may take more than one drug, and there are 25–30 different drugs under investigation.
The feasible maximum sample size is about 100 mothers. Given these numbers, I think the only realistic option may be to estimate a global correlation coefficient that accounts for the multilevel data structure.
What would you do in this situation? What approach to calculate an adequate sample size?
Could one possibility be to increase the sample size using a design effect, as we usually do in multicenter prevalence studies?
Thank you in advance for your help.
Gianfranco
I am having issues with a sample size calculation and would appreciate your advice.
Here is the scenario: I have mothers who are taking medications for an encephalopathy and who are also breastfeeding.
These drugs can be transferred into breast milk and then into the infants’ bloodstream.
I need to quantify the correlation between the amount of drug in breast milk and the amount of drug in the infants’ blood.
If we assumed that each mother was taking only one drug, and that only one drug was being studied, the problem would be straightforward (e.g., assuming a a priori Pearson correlation coefficient of 0.4, we would need roughly 40–50 mother–infant pairs).
However, in this study each mother may take more than one drug, and there are 25–30 different drugs under investigation.
The feasible maximum sample size is about 100 mothers. Given these numbers, I think the only realistic option may be to estimate a global correlation coefficient that accounts for the multilevel data structure.
What would you do in this situation? What approach to calculate an adequate sample size?
Could one possibility be to increase the sample size using a design effect, as we usually do in multicenter prevalence studies?
Thank you in advance for your help.
Gianfranco

Comment