Dear statalist users,
In a study of social mobility in Brazil I am estimating returns to complete tertiary education (versus lower level) on children's income conditional on class origin, race, and cohort. I use a Generalized Linear Model (Gamma distribution; logarithmic link function) with simultaneous interactions between class origin, race, education and cohort and controls only for factors prior to the child's entry into the labor market. It should be considered that the Generalized Linear Model with a logarithmic link function has the advantage of generating predicted values at real levels before applying the link function, that is, the model avoids any transformation bias (Baum, 2021).
Estimates for the most recent ten-year cohort and origin at the social top (on the dimensions of capital, authority, and expert knowledge) on proportional returns to higher education (margins command eydx option) show this unexpected result: the white group reveals a disadvantage relative to the brown & black group of the order of -0.363 in log or -30.4% (p=0.028).
Estimates with the same model and specifications, but in terms of absolute differences (dydx option of margins) yield this seemingly contradictory result: the white group has a significant disadvantage of -980 Reais (Brazilian currency), which represents a difference of -27.9%. However, this difference is uncertain in the statistical significance criterion (p=0.137), even though the confidence interval is strongly inclined towards the negative direction. I am aware of the limitations of the p-value criterion that Clyde Schechter strongly emphasizes.
Note that the discrepancy between the estimates concerns a second-order difference, that is, a (racial) difference between differences (in returns to education) using the mlincom option for Stata.
Interpretation: The estimates of proportional and absolute differences do not differ in the size and direction of the effect, that is, the discrepancies between the estimates basically concern a problem of precision and statistical significance. It seems revealing that the problem with statistical significance only occurs in the post-estimation when bringing the predicted values to the absolute scale in Reais (currency in Brazil) and estimating the difference and standard error for this scale since both estimates were made with the same model with the same specifications. The discrepancies found in the results would arise purely from the fact that one measure expresses a relative change and the other an absolute change. As long as the effect is strong and there is no shortage of cases, it seems reasonable to assume that incomes in absolute values show greater variability (particularly at higher income levels), which may affect the standard error and statistical significance. This variability problem could be aggravated by an estimate of the difference between differences with the respective levels of variability underlying the differences.
Would this be the main conclusion to be emphasized? That's what I have concluded.
In substantive and statistical terms, the striking evidence is the size and direction of the effect, which is convergent in both estimates and shows that in origin at the social top, the returns to higher education are clearly unfavorable to the white group in the most recent cohort.
Would this alternative conclusion be more appropriate?
Both results are valid in what they are expressing and must be considered in the light of the characteristic of each measure (absolute and relative) and the type of manifestation of the phenomenon under investigation.
Would another interpretation be more adequate?
Could this discrepancy have another source? What would this source be and what is its implication?
Given the relevance and substantive implication of the evidence and the issue, all comments are welcome!
In a study of social mobility in Brazil I am estimating returns to complete tertiary education (versus lower level) on children's income conditional on class origin, race, and cohort. I use a Generalized Linear Model (Gamma distribution; logarithmic link function) with simultaneous interactions between class origin, race, education and cohort and controls only for factors prior to the child's entry into the labor market. It should be considered that the Generalized Linear Model with a logarithmic link function has the advantage of generating predicted values at real levels before applying the link function, that is, the model avoids any transformation bias (Baum, 2021).
Estimates for the most recent ten-year cohort and origin at the social top (on the dimensions of capital, authority, and expert knowledge) on proportional returns to higher education (margins command eydx option) show this unexpected result: the white group reveals a disadvantage relative to the brown & black group of the order of -0.363 in log or -30.4% (p=0.028).
Estimates with the same model and specifications, but in terms of absolute differences (dydx option of margins) yield this seemingly contradictory result: the white group has a significant disadvantage of -980 Reais (Brazilian currency), which represents a difference of -27.9%. However, this difference is uncertain in the statistical significance criterion (p=0.137), even though the confidence interval is strongly inclined towards the negative direction. I am aware of the limitations of the p-value criterion that Clyde Schechter strongly emphasizes.
Note that the discrepancy between the estimates concerns a second-order difference, that is, a (racial) difference between differences (in returns to education) using the mlincom option for Stata.
Interpretation: The estimates of proportional and absolute differences do not differ in the size and direction of the effect, that is, the discrepancies between the estimates basically concern a problem of precision and statistical significance. It seems revealing that the problem with statistical significance only occurs in the post-estimation when bringing the predicted values to the absolute scale in Reais (currency in Brazil) and estimating the difference and standard error for this scale since both estimates were made with the same model with the same specifications. The discrepancies found in the results would arise purely from the fact that one measure expresses a relative change and the other an absolute change. As long as the effect is strong and there is no shortage of cases, it seems reasonable to assume that incomes in absolute values show greater variability (particularly at higher income levels), which may affect the standard error and statistical significance. This variability problem could be aggravated by an estimate of the difference between differences with the respective levels of variability underlying the differences.
Would this be the main conclusion to be emphasized? That's what I have concluded.
In substantive and statistical terms, the striking evidence is the size and direction of the effect, which is convergent in both estimates and shows that in origin at the social top, the returns to higher education are clearly unfavorable to the white group in the most recent cohort.
Would this alternative conclusion be more appropriate?
Both results are valid in what they are expressing and must be considered in the light of the characteristic of each measure (absolute and relative) and the type of manifestation of the phenomenon under investigation.
Would another interpretation be more adequate?
Could this discrepancy have another source? What would this source be and what is its implication?
Given the relevance and substantive implication of the evidence and the issue, all comments are welcome!

Comment